Paper Feed

Curated ML papers with one-line takes on why they matter.

News & Events

Recent happenings in AI research.

2026-04-23ReleaseLATEST

DeepSeek-V4 Officially Open-Sourced — 1.6T MoE, 1M Context

DeepSeek-V4 is officially live and open-sourced. V4-Pro: 1.6T total / 49B active params, 1M context, 80.6% SWE Verified, 93.5 LiveCodeBench, Codeforces rating 3206. V4-Flash: 284B total / 13B active, same 1M context. Both support Expert Mode (Think) and Instant Mode (Non-Think). Pre-trained on 32T+ tokens with Muon optimizer. Weights on HuggingFace, API available, tech report released.

DeepSeek Official ↗
2026-04-21Release

GPT Image 2.0 — Reasoning-First Image Generation

OpenAI launched gpt-image-2 (ChatGPT Images 2.0) with integrated o-series reasoning — the model plans and reasons before generating. Supports up to 8 coherent images per prompt, accurate multilingual text rendering (Chinese, Japanese, Korean), and 2K resolution. Hit #1 on Image Arena within 12 hours by a +242 point margin.

OpenAI Blog ↗
2026-04-04Industry

Claude Code Source Leaked via npm Source Maps

Anthropic's Claude Code CLI source code was inadvertently exposed via npm source maps, revealing 1,884 TypeScript files across 36 folders. The leak exposed internal feature flags, unreleased agent modes (ultraplan, kairos-proactive), and architecture details. Anthropic has since patched the package.

ccleaks.com Analysis ↗
2026-04-03Industry

3 Security Flaws in Claude Code Allow Remote Code Execution

Check Point Research identified three vulnerabilities (CVE-2025-59536, CVE-2026-21852) in Claude Code that allow attackers to run arbitrary code and steal API keys via malicious repositories.

Check Point Research ↗
2026-02-17Release

Claude Sonnet 4.6 Released

Anthropic released Claude Sonnet 4.6, delivering frontier performance across coding, agents, and professional work at scale.

Anthropic News ↗

179 papers

Sep 2026

⭐

A Ticket from Marginals to Joints: Coupled-Noise Distillation for One-Step Block Generation in Diffusion Language Models

Lin Yao · 2026-09

CONDOR uses coupled-noise distillation to train a student to generate a full block in one forward pass, without a target-side encoder or autoregressive teacher.

⭐

Language Models Can Control Their Own Attention

Namgyu Ho, Huzama Ahmad, Woosung Koh, Se-Young Yun, Tal Schuster, and Cicero Nogueira dos Santos · 2026-09

Declarative Attention has the model emit global, focus, and local attention scopes; in zero-shot evaluation on 15 long-context tasks, the authors report fewer attended tokens with modest accuracy drops.

⭐

BASIS: Batchwise Advantage Estimation from Single-Rollout Information Sharing for LLM Reasoning

Shijin Gong, Erhan Xu, Kai Ye, Giulia Livieri, Francesco Quinzan, and Chengchun Shi · 2026-09

BASIS shares information across prompts in a batch to estimate advantages from one rollout each; the authors report 69% lower MSE than REINFORCE++ in their experiments.

·

Predict, Don't Iterate: Efficient Adaptive-Length Infilling for Diffusion Language Models

Haobo Xu et al. · 2026-09

PILL predicts infill length from one mask-token hidden state, enabling variable-length infilling in fixed-length diffusion LMs in two passes; the authors report a 4.8-point code pass-rate gain and 1.82× speedup. EMNLP 2026 main conference.

·

Routing by Reasoning Need: Trajectory-Aware Decoding Control for Diffusion Vision-Language Models

Yixiang Liu, Zhongxing Xu, Zhonghua Wang, and Xiaoying Tang · 2026-09

A training-free controller routes LLaDA-V queries among early, baseline, and reasoning-supportive decoding trajectories; the authors report improved robustness to task-dependent reasoning needs. Findings of EMNLP 2026.

·

Coupled Calibration and Learning: Mitigating Teacher Bias in LLM Distillation without Target-Domain Reward Feedback

Haichen Hu, Yuheng Zhang, and David Simchi-Levi · 2026-09

CCL calibrates a teacher using source-domain rewards before training a target-domain student. Under its stated framework, the paper proves that the student's KL divergence to an oracle in the student class converges to zero.

⭐

NCP-ArchPreview Technical Report: Moving towards Latent Space Language Models through Next Concept Prediction

The Intern-NCP Team · 2026-09

Jointly trains next-token and next-concept prediction: a product-quantized latent vocabulary represents multi-token concepts while preserving token-level autoregressive generation.

⭐

Unlocking Lossless Speedups in LLMs via Discrete Diffusion

Subham Sekhar Sahoo et al. · 2026-09

The authors augment autoregressive LMs with lightweight diffusion weights and Ψ-Spec samplers for parallel decoding; they report up to 3× throughput over the base AR model.

⭐

Does Deeper Reasoning Compromise Alignment? Revealing and Mitigating Alignment Collapse in Large Reasoning Models

Yu-Hang Wu et al. · 2026-09

The paper reports that deeper reasoning can raise its proposed Alignment Loss Rate under perturbation, and introduces Reasoning Residual Alignment as a mitigation.

·

PELM: Power Efficient On-Device LLM Inference with Speculative Decoding and Dynamic Voltage Frequency Scaling

Weisi Yang and Stephen Xia · 2026-09

Combines speculative decoding, variable verification depth, and DVFS for on-device inference; the authors report up to 23.1% speedup and 52.4% lower energy use at comparable task performance.

·

SpatialBlock: Enhancing Spatial Intelligence in LVLMs via Synthetic Block-Stacking Problem

Soohyun Ryu et al. · 2026-09

SpatialBlock-15k supplies 15,000 synthetic block-stacking problems for 3D-to-2D projection, viewpoint transformation, and structural-composition reasoning; the authors report gains on real-world spatial tasks.

⭐

Semantic-Space Exploration and Exploitation in RLVR for LLM Reasoning

Fanding Huang et al. · 2026-09

Studies how RLVR shapes semantic exploration in LLM reasoning — proposes strategies to improve coverage of valid reasoning paths.

⭐

Efficient Pre-Training with Token Superposition

Bowen Peng, Théo Gigant, and Jeffrey Quesnelle · 2026-09

Token-Superposition Training combines superposition and recovery phases; the authors report up to 2.5× lower total pretraining time at equal loss for a 10B A1B model.

·

Integrated and Cross-Architecture Interpretation of LLM Reasoning

Leonardo Matthew Yauw, Wei-Bin Kou, and Yujiu Yang · 2026-09

Unified interpretability framework across LLM architectures traces and compares internal reasoning representations — enables systematic mechanistic understanding.

·

Encoder-Decoder Diffusion Language Models for Efficient Training and Inference

Marianne Arriola et al. · 2026-09

E2D2 is an encoder-decoder diffusion language model framework; the authors report improvements in training and inference efficiency.

Aug 2026

⭐

LLaDA MoE v2: Scaling Mixture-of-Experts Diffusion Language Models

Fengqi Zhu et al. · 2026-08

First systematic MoE scaling study for diffusion LMs — 30B-A3B on 23.5T tokens, revealing that AR scaling laws don't transfer directly to diffusion models.

⭐

SelFusion: Self-distillation for Diffusion Language Models

Hyeongsoo Lim et al. · 2026-08

ACL 2026: bidirectional self-distillation between hard/easy masking schedules fixes KD failure modes in diffusion LMs — trained students routinely outperform LLM teachers.

⭐

MUX: Continuous Reasoning via Multiplexed Tokens

Ayhan Suleymanzade et al. · 2026-08

Each MUX latent token is a weighted superposition of multiple reasoning steps — enables parallel exploration without beam-search overhead, beating latent baselines across 32 settings on 4 models.

·

xPress: Parallel Refinement for Diffusion Drafters in Speculative Decoding

Anonymous et al. · 2026-08

Lightweight causal refiner restores joint distribution from per-position marginals in block-diffusion drafters — improves acceptance rate with one parallel pass, no extra model needed.

·

Dependency-Aware Revocable Decoding for Efficient Diffusion LLM Inference

Anonymous et al. · 2026-08

Models token dependency structure to selectively revoke committed positions during decoding — reduces quality loss from premature commits in masked diffusion LMs.