
From spacetime patches and Diffusion Transformers to open STDiT stacks. Research lineage only: what ‘world simulator’ claims mean, what Open-Sora actually ships, and where eval still lies.

Lewis et al. Retrieval-Augmented Generation and the dense-retrieval lineage. When “Cultural LLM” should mean corpus + retrieve, not only weights.

From human preference models and PPO to Direct Preference Optimization. What operators actually buy when a vendor says ‘aligned,’ and which failure modes the papers already named.

Nanjing University / Tencent Youtu / CASIA on masked discrete diffusion any-to-any. Text, speech, and images without an AR MLLM backbone.

From sparsely-gated MoE to Switch Transformer and Mixtral. How sparse experts scale parameters without paying dense compute on every token.

Adobe/UCLA/Georgia Tech on multimodal diffusion reasoning. Unified SFT+RL, answer-forcing, tree search, and the numbers that moved.

Dao et al. on IO-aware exact attention. Why Transformers got memory- and bandwidth-efficient without approximating the math.

Tokenization, benchmarks, and data realities for Creoles and other low-resource languages — with JamPatoisNLI, CreoleVal, and NLLB as anchors for JA/TT speech-text work.

Whisper’s weak-supervision lineage, Jamaican Patois fine-tuning numbers, Caribbean emergency triage work, and why LibriSpeech wins do not transfer to WhatsApp voice notes.

Vaswani et al. 2017 in plain terms. What the Transformer claimed, what changed after it, and who should still care in 2026.