Overview
An experimental investigation into discrete diffusion language modeling (non-autoregressive generation) as an alternative to traditional causal next-token prediction.
The system trains a bidirectional Transformer Denoiser to iteratively refine fully masked or noisy token sequences into coherent natural language text under arbitrary fixed-token constraints.
Technical Details
1. Architecture & Denoising Objectives
- Implemented a ~253M parameter bidirectional Transformer denoiser in PyTorch.
- Built time-step embedding schedules and discrete transition matrices for continuous-to-discrete noise schedules.
- Formulated categorical diffusion and masked language modeling objectives (MDLM) over 1.0B English tokens (SlimPajama corpus).
2. Controllable Text Infilling & Fixed-Token Sampling
- Implemented conditional inference algorithms allowing users to pin tokens at arbitrary sequence indices (prefix, suffix, or infill).
- The diffusion sampling loop preserves fixed positions during reverse denoising steps, generating syntactically and semantically coherent text connecting boundary constraints.
- Benchmarked offline inference performance and perplexity metrics across Apple Silicon (MPS) and CUDA environments.
3. Findings
- Token infill diffusers are still a challenge on small token count and limited hardware.
- This will be revisited in the the future at some point.