<?xml version="1.0" encoding="UTF-8"?>
<oai_dc:dc xmlns:oai_dc="http://www.openarchives.org/OAI/2.0/oai_dc/"
           xmlns:dc="http://purl.org/dc/elements/1.1/"
           xmlns:xsi="http://www.w3.org/2001/XMLSchema-instance"
           xsi:schemaLocation="http://www.openarchives.org/OAI/2.0/oai_dc/ http://www.openarchives.org/OAI/2.0/oai_dc.xsd">
  <dc:title>Where Quality Breaks in Compressed Short-Text Generation: Staged Bottleneck Localization</dc:title>
  <dc:creator>Alexey Gavrilov</dc:creator>
  <dc:creator>Alan-Barsag Gazzaev</dc:creator>
  <dc:creator>Sergey Muravyov</dc:creator>
  <dc:subject>compressed text generation</dc:subject>
  <dc:subject>discrete latent diffusion</dc:subject>
  <dc:subject>hierarchical VQ-VAE-2</dc:subject>
  <dc:subject>bottleneck localization</dc:subject>
  <dc:subject>TinyStories</dc:subject>
  <dc:subject>masked diffusion language modeling</dc:subject>
  <dc:description>Compressed short-text generators can fail in two different places: the codec may discard information before generation starts, or the latent generator may produce weak codes. Without separating these failure modes, researchers can spend compute improving the wrong component. We study this problem in a controlled 64-to-16 TinyStories case study built from a hierarchical VQ-VAE-2 codec and a masked discrete diffusion generator (MDLM). We use a staged validation protocol that separates codec reconstruction fidelity, latent generation quality, and auxiliary latent diagnostics under one shared external GPT-2 scorer, while reporting complementary semantic metrics for the geometry study. In the tested configuration, codec reconstruction alone raises median external perplexity from 15.17 to 27.36 (+80.4%) and p95 from 25.10 to 98.91 (+294.1%), showing that the dominant quality loss appears before latent generation begins. Under the same scorer, code-space MDLM remains materially stronger than token-space diffusion, reducing mean, median, and p95 by 32.9%, 30.9%, and 36.6%, respectively. Geometry-aware regularization improves local latent proxies but does not improve decoded-text metrics in the available runs. The contribution is methodological rather than algorithmic: the paper presents a reusable staged diagnosis for one concrete pipeline and shows that, in this setting, codec fidelity rather than latent denoising sets the practical quality ceiling.</dc:description>
  <dc:publisher>IEEE</dc:publisher>
  <dc:date>2026-04-28</dc:date>
  <dc:type>Text</dc:type>
  <dc:format>text/html</dc:format>
  <dc:format>application/pdf</dc:format>
  <dc:identifier>https://doi.org/10.23919/FRUCT70069.2026.11506553</dc:identifier>
  <dc:identifier>https://aogavrilov.com/publications/where-quality-breaks/</dc:identifier>
  <dc:source>2026 39th Conference of Open Innovations Association (FRUCT)</dc:source>
  <dc:language>en</dc:language>
  <dc:relation>https://aogavrilov.com/publications/where-quality-breaks/paper.pdf</dc:relation>
  <dc:relation>https://aogavrilov.com/ru/publications/where-quality-breaks/</dc:relation>
  <dc:relation>https://aogavrilov.com/zh/publications/where-quality-breaks/</dc:relation>
  <dc:rights>© 2026 IEEE. Personal use of this material is permitted. Permission from IEEE must be obtained for all other uses, in any current or future media, including reprinting or republishing this material for advertising or promotional purposes, creating new collective works, for resale or redistribution to servers or lists, or reuse of any copyrighted component of this work in other works.</dc:rights>
  <dc:rights>https://journals.ieeeauthorcenter.ieee.org/become-an-ieee-journal-author/publishing-ethics/guidelines-and-policies/post-publication-policies/</dc:rights>
</oai_dc:dc>
