CV
A one-page PDF is available above; the full version follows below.
Basics
| Name | Kaustubh Ponkshe |
| Label | PhD Student in Computer Science, EPFL |
| kaustubh.ponkshe@epfl.ch | |
| Url | https://KaustubhP11.github.io |
| Summary | PhD student at EPFL (MLO lab, advised by Prof. Martin Jaggi) and contributor to Apertus. I study how language models learn — pretraining and midtraining dynamics, memory and knowledge, and the efficient and reliable adaptation of foundation models. |
Education
-
2025.09 - Present Lausanne, Switzerland
Ph.D. (EDIC)
École Polytechnique Fédérale de Lausanne (EPFL)
Computer Science (Machine Learning & Optimization Lab, advised by Prof. Martin Jaggi)
-
2019.07 - 2024.07 Mumbai, India
Dual Degree (B.Tech + M.Tech)
Indian Institute of Technology Bombay (IITB)
Electrical Engineering (B.Tech) and Artificial Intelligence / Machine Learning (M.Tech)
Work
-
2025.06 - 2025.08 London / Remote
Founder in Residence
Entrepreneurs First
Selected for a curated cohort of founders building globally impactful companies.
- Explored protein-language models to accelerate discovery of novel proteins for sustainable food synthesis
-
2024.06 - 2025.06 Abu Dhabi, UAE
Researcher
Mohamed bin Zayed University of Artificial Intelligence (MBZUAI)
LLM efficiency, AI safety and security, post-training, and large-scale distributed learning, with Prof. Praneeth Vepakomma (CERT-Lab).
- Made LLM fine-tuning more efficient, reducing parameter requirements by up to 90x without loss of performance
- Designed update schemes for distributed adaptation, cutting communication overhead by up to 230x
- Studied the impact of fine-tuning on safety and alignment, exposing limits of subspace-based defenses
-
2022.08 - 2023.04 Bengaluru, India
Research Collaborator
Adobe Research
Structure-aware pre-training for scientific document understanding.
- Built a structure-aware corpus of 100K arXiv documents and pre-trained a Longformer with global tokens
- Achieved an 18% gain over sparse-local baselines on SciREX information extraction
Publications
-
2025.09.01 Apertus: Democratizing Open and Compliant LLMs for Global Language Environments
Project Apertus (Swiss AI Initiative)
-
2025.05.01 Safety Subspaces are Not Distinct: A Fine-Tuning Case Study
ICLR 2026; INTERPLAY @ CoLM 2025
-
2025.05.01 ABBA: Highly Expressive Hadamard Product Adaptation for Large Language Models
ICLR 2026; ES-FoMo @ ICML 2025 (Spotlight)
-
2025.02.01 Fed-SB: A Silver Bullet for Extreme Communication Efficiency and Performance in (Private) Federated LoRA Fine-Tuning
TMLR (J2C Certification — top 10%)
-
2025.02.01 TokenSwap: A Lightweight Method to Disrupt Memorized Sequences in LLMs
NeurIPS 2025 (Spotlight)
-
2024.11.01 -
2024.10.01 FedEx-LoRA: Exact Aggregation for Federated and Efficient Fine-Tuning of Large Language Models
ACL 2025 (Oral, Main Conference)
Awards
- 2025.09.01
- 2024.05.01
- 2019.06.01
- 2017.12.01
Skills
| Research | |
| LLM Pretraining & Midtraining | |
| Post-Training & Data Mixtures | |
| Memorization & Generalization | |
| Parameter-Efficient Fine-Tuning | |
| Federated Learning | |
| AI Safety |
| Languages & Libraries | |
| Python | |
| PyTorch | |
| Megatron-LM | |
| Nanotron | |
| DeepSpeed | |
| Transformers | |
| PEFT | |
| TRL | |
| vLLM | |
| Accelerate |
Languages
| English | |
| Fluent |
| Hindi | |
| Native |
| Marathi | |
| Native |
| French | |
| Basic (learning) |
Interests
| Outside Research | |
| Traveling (24 countries and counting) | |
| Tennis |