Research that stays inspectable.
I study how representations, compression, and controllability shape the quality of generated text and code.
Explore controllable code editingFeatured research
Diagnostic methodology
Where Quality Breaks in Compressed Short-Text Generation: Staged Bottleneck Localization
Separating codec failures from latent-generation failures in compressed short-text generation.
Explore codec bottleneck diagnosisOther research
Browse all publicationsResearch directions
Codec or generator failure?
Separate reconstruction loss from latent-generation loss before spending compute on the wrong stage.
Diagnose the failing stageLocalized and predictable code editing
Prevent full-function rewrites by comparing selective regeneration, constrained generation, and AI-assisted refactoring under an explicit preservation contract.
Choose a code-editing control surfaceCode-space or token-space MDLM?
Compare both under one decoded-text scorer while accounting for the quality ceiling imposed by a lossy codec.
Compare the reported stagesBrowse 5 focused research questions
- How should code-space and token-space masked diffusion language modeling be compared?
- How can a generative model modify code without rewriting the entire function?
- What does constrained code generation guarantee in a software-engineering workflow?
- Which methods and evidence matter for AI-assisted behavior-preserving refactoring?
- What makes code generation predictable rather than merely controllable?