Ende Shen

Independent researcher · robot learning

prof_pic.jpg

I’m an independent researcher in San Francisco working on robot learning. I study what happens to vision-language-action (VLA) models when they keep learning after deployment: which behaviors survive fine-tuning, which are silently lost, and what that means for safety. I run experiments on π0.5, both in simulation and on a low-cost real arm.

My background is in verifiable ML. As Staff Cryptographer at Modulus Labs, I built Remainder, a zero-knowledge proving system for ML inference, and at Tools for Humanity I implemented MPC proofs for a system serving ~22M users. I also founded Tegore (YC S25), a real-time voice tutor for K-12 math. The thread through all of it: how do we know what a system will do before it acts?

I did my B.S. and M.S. in Computer Science at Stanford, where I worked on faithful text generation in Tatsu Hashimoto’s group.

On my mind

  • Can we tell from a VLA’s internals what fine-tuning will erase before it happens?
  • Do the safety behaviors trained into robot foundation models survive the routine fine-tuning that happens after deployment?
  • Should self-improvement methods be judged by a frozen benchmark score, or by how fast they improve under a fixed budget?

selected publications

  1. Data Source Decides Forgetting: Self-Generated Rollouts vs. Human Demonstrations for a Flow-Matching VLA
    Ende Shen
    2026
    Under review
  2. Learned Collision Avoidance in VLAs Does Not Survive Routine Fine-Tuning
    Ende Shen
    2026
    Under review
  3. ACL
    TempLM: Distilling Language Models into Template-Based Generators
    Tianyi Zhang, Mina Lee, Xiang Lisa Li, Ende Shen, and Tatsunori Hashimoto
    In Findings of the Association for Computational Linguistics: ACL 2023, 2023