CV

A one-page PDF is available above; the full version follows below.

Basics

Name Kaustubh Ponkshe
Label PhD Student in Computer Science, EPFL
Email kaustubh.ponkshe@epfl.ch
Url https://KaustubhP11.github.io
Summary PhD student at EPFL (MLO lab, advised by Prof. Martin Jaggi) and contributor to Apertus. I study how language models learn — pretraining and midtraining dynamics, memory and knowledge, and the efficient and reliable adaptation of foundation models.

Education

  • 2025.09 - Present

    Lausanne, Switzerland

    Ph.D. (EDIC)
    École Polytechnique Fédérale de Lausanne (EPFL)
    Computer Science (Machine Learning & Optimization Lab, advised by Prof. Martin Jaggi)
  • 2019.07 - 2024.07

    Mumbai, India

    Dual Degree (B.Tech + M.Tech)
    Indian Institute of Technology Bombay (IITB)
    Electrical Engineering (B.Tech) and Artificial Intelligence / Machine Learning (M.Tech)

Work

  • 2025.06 - 2025.08

    London / Remote

    Founder in Residence
    Entrepreneurs First
    Selected for a curated cohort of founders building globally impactful companies.
    • Explored protein-language models to accelerate discovery of novel proteins for sustainable food synthesis
  • 2024.06 - 2025.06

    Abu Dhabi, UAE

    Researcher
    Mohamed bin Zayed University of Artificial Intelligence (MBZUAI)
    LLM efficiency, AI safety and security, post-training, and large-scale distributed learning, with Prof. Praneeth Vepakomma (CERT-Lab).
    • Made LLM fine-tuning more efficient, reducing parameter requirements by up to 90x without loss of performance
    • Designed update schemes for distributed adaptation, cutting communication overhead by up to 230x
    • Studied the impact of fine-tuning on safety and alignment, exposing limits of subspace-based defenses
  • 2022.08 - 2023.04

    Bengaluru, India

    Research Collaborator
    Adobe Research
    Structure-aware pre-training for scientific document understanding.
    • Built a structure-aware corpus of 100K arXiv documents and pre-trained a Longformer with global tokens
    • Achieved an 18% gain over sparse-local baselines on SciREX information extraction

Awards

Skills

Research
LLM Pretraining & Midtraining
Post-Training & Data Mixtures
Memorization & Generalization
Parameter-Efficient Fine-Tuning
Federated Learning
AI Safety
Languages & Libraries
Python
PyTorch
Megatron-LM
Nanotron
DeepSpeed
Transformers
PEFT
TRL
vLLM
Accelerate

Languages

English
Fluent
Hindi
Native
Marathi
Native
French
Basic (learning)

Interests

Outside Research
Traveling (24 countries and counting)
Tennis