Broteína: Self-Distillation Unlocks Few-Step Protein Design
Aligning the teacher with its useful low-temperature design distribution before compression lets few-step students surpass the 400-step base model's distinct designable yield.
I am a master's student at MIT interested in building more efficient machine learning models: architectures that do more with less compute, less memory, and less friction between an idea and a working implementation.
At MIT, I have been lucky to work with Prof. Daniela Rus and Prof. Song Han on efficient ML across different layers of the stack, from sparse architecture design, hardware-aware kernel design, distillation, and data-level methods. I am currently tinkering with data and training for LLM pre-training and post-training.
A long time ago, I was an International Olympiad in Informatics (IOI) Gold Medalist, surpassed 2900 on Codeforces, and spent a lot of time writing programming problems for CodeChef and other platforms. I also worked on educational projects, including contributing to a freely available book on Computational Graph Theory for Olympiad (GTOI) among other resources, which shaped how I think about access, teaching, and the joy of making difficult ideas easier for other people to enter.
Aligning the teacher with its useful low-temperature design distribution before compression lets few-step students surpass the 400-step base model's distinct designable yield.
A scaling story for oscillator state-space layers: block-sparse projection heads control dense coupling, while IO-aware FlashDOSS kernel fuses projection, scan, and projection-back work to avoid expensive state-domain traffic.
Guangxuan Xiao*, Junxian Guo*, Kasra Mazaheri, Song Han
Flash MoBA studies how mixture-of-block attention can route long-context computation sparsely while preserving quality, connecting routing accuracy, block size, local key aggregation, and efficient CUDA execution.
Kasra Mazaheri, Mohammed Ehab
Preprint available to share
Broteína aligns a protein backbone model's native score field with its useful low-temperature design distribution before compression, enabling few-step students that surpass the base model's distinct designable yield.
[blog]
Kasra Mazaheri, Jared Boyer, T. Konstantin Rusch, Daniela Rus
Preprint available to share
FlashLinOSS shows that block-sparse oscillator SSMs can outperform denser variants with fewer projection parameters, while IO-aware fused kernels reduce runtime by up to 7.8x and peak memory by about 3x.
Cambridge, MA
M.S. in Computer Science and Artificial Intelligence, expected 2026
B.S. in Computer Science and Engineering, minors in Mathematics and Music Technology, 2025
GPA: 5.0/5.0. Invited member of Phi Beta Kappa.
Quantitative Research, Jun 2025 - Aug 2025
Algorithm Engineering, May 2024 - Aug 2024
Software Engineering, Jun 2023 - Aug 2023