A NeurIPS 2026 Workshop · Sydney, Australia

BeNTo: Beyond Next‑Token Prediction —
Diffusion & Flow Models for Next‑Generation Decoding

Beyond next-token prediction — exploring the theory, algorithms, and applications of discrete diffusion and flow models for parallel, non-causal generation.

Submissions dueSep 5, 2026
NotificationSep 28, 2026
Workshop · single-dayDec 11 or 12, 2026

01 Overview

Rethinking the order of generation.

Autoregressive models produce their output one element at a time, in a fixed order. The recipe is remarkably general, and it has carried the field a long way. However, it also ties the order of computation to the order of the result, leaving parallelism, revision, and control hard to reach. Discrete diffusion and flow models take a different route. They generate by iterative denoising: refining many positions at once, and revisiting earlier choices as a sample takes shape.

That shift brings a different set of computational properties — non-causal, parallel, and controllable — along with open questions spanning mathematics, algorithms, and engineering, and a growing body of work that combines the two paradigms rather than choosing between them. This workshop joins two conversations: the theory and algorithms that make discrete generative models work, and the applications and systems that put them to use.

Track 01

Depth — Theories & Algorithms

Discrete diffusion and flow models at the intersection of generative modeling, optimal transport, and stochastic optimal control.

Generative Modeling
Discrete diffusion & flow models, optimal transport, optimal control, Schrödinger bridges.
Probabilistic Inference
Discrete diffusion samplers, adjoint-based samplers, advanced MCMC, variational inference.
Track 02

Breadth — Applications & Systems

Novel applications and scalable systems that exploit non-causal, parallel generation.

Applications
AI for science, multimodal generation, reward alignment, inverse problems, benchmarks & datasets.
Systems & Empirical Analysis
Foundation models, large-scale training/inference, network architectures.

02 Scope

Topics of interest.

This workshop considers, but is not limited to, the following topics.

IDiscrete & continuous diffusion/flow models
IIDiffusion & flow for language modeling
IIIAI for science
IVMultimodal generation
VReward alignment
VIDiffusion & flow samplers
VIISchrödinger bridges
VIIIOptimal & stochastic control
IXOptimal transport
XFoundation models & large-scale training/inference
XISystems & hardware for diffusion and flow models
XIIBenchmarks & datasets

03 Call for Papers

Share your work.

What to submit

We invite submissions on any topic within the workshop's scope, across both tracks, from theory and algorithms to applications and systems. All submissions are handled through OpenReview.

  • Format. Use the NeurIPS 2026 LaTeX template and anonymize author information.
  • Page limit. Submissions may be either 4 or 8 pages of main text. References and appendices do not count toward the page limit.
  • Mutual-review policy. At least one qualified author per submission serves as a reviewer.
  • Travel grants for students and early-career researchers with accepted papers.

Presentation

Accepted papers are presented as posters, with a selection of contributed orals. A Best Paper Award recognizes outstanding work.

04 Important Dates

Mark your calendar.

Jul 20, 2026Call for Papers & submissions open OpenReview
Sep 5, 2026Paper submission deadline
Sep 6 – 20, 2026Peer review period
Sep 21 – 25, 2026Meta-review & discussion
Sep 28, 2026Accept / reject notification
Oct 20, 2026Camera-ready deadline tentative
Dec 11 or 12, 2026Workshop @ NeurIPS 2026 single-day; exact day TBD

All deadlines are 23:59 AoE. Dates are indicative and subject to change.

05 Schedule

A one-day program.

A full day of invited talks, contributed orals, two poster sessions, and a panel. Speaker slot assignments are being finalized.

08:50–09:00Opening remarksWelcome
09:00–09:30Invited Talk 1Keynote
09:30–10:00Invited Talk 2Keynote
10:00–11:00Poster session 1 & breakPosters
11:00–11:30Invited Talk 3Keynote
11:30–12:00Invited Talk 4Keynote
12:00–12:30Oral session 1 (3 papers)Orals
12:30–13:30LunchBreak
13:30–14:30Panel discussionPanel
14:30–15:00Invited Talk 5Keynote
15:00–15:30Invited Talk 6Keynote
15:30–16:00Oral session 2 (3 papers)Orals
16:00–17:00Poster session 2 & socialPosters
17:00–17:10Awards & closingBest Paper

Tentative program · times shown in the local conference time zone. Invited-talk slots will be mapped to speakers closer to the event.


06 Invited Speakers

Six voices shaping the field.

All six speakers are confirmed. Talk titles will be announced closer to the workshop.

SEPortrait of Stefano Ermon
Stanford University

Associate Professor of Computer Science at Stanford, where his group advances generative modeling, probabilistic inference, and diffusion methods. His recent work pushes diffusion language models toward production-scale text generation.

MUPortrait of Masatoshi Uehara

Member of the technical staff at OpenAI, working on next-generation language models for scientific discovery. His research spans reinforcement learning, test-time steering, and post-training of diffusion and discrete generative models, with applications to protein and drug design.

JSPortrait of Jiaxin Shi
Meta Superintelligence Labs

Research scientist at Meta Superintelligence Labs, previously at Google DeepMind. He works on generative modeling and probabilistic inference — masked diffusion, gradient estimation, and sampling — and on sequence models that encode long-range dependencies efficiently.

GLPortrait of Guan-Horng Liu
Meta Superintelligence Labs

Research scientist at Meta Superintelligence Labs, where he builds generative models with scientific structure. His work on Schrödinger bridges and adjoint-based diffusion samplers brings optimal-transport and control priors to diffusion, with applications from image restoration to protein folding and particle physics.

AGPortrait of Aditya Grover
UCLA · Inception

Assistant Professor of Computer Science at UCLA, where he leads the Machine Intelligence (MINT) group, and co-founder and CTO of Inception, which builds ultra-fast parallel diffusion LLMs. His research sits at the intersection of generative models, reinforcement learning, and scientific discovery.

AVPortrait of Arash Vahdat
NVIDIA Research

Research Director at NVIDIA Research, leading the Fundamental Generative AI Research (GenAIR) team. His work spans diffusion and flow models, discrete generative models, and accelerated sampling, applied to images, video, weather, proteins, and molecules.


07 Organizers

Brought to you by the organizing team.

General enquiries: bento-neurips@googlegroups.com

NDPortrait of Nolan Dey
Cerebras Systems

Research Scientist at Cerebras Systems working on large-scale training, scaling laws, and efficient deep learning.

08 Sponsors & Support

Support that widens the door.

Sponsorship helps us fund travel grants and broaden participation for students and early-career researchers. We gratefully acknowledge the organizations supporting this workshop.

Ready to submit?

Join us at NeurIPS 2026 to push generation beyond next-token prediction. Submissions open July 20 and close September 5, 2026.

Submit on OpenReview

Submissions open Jul 20, 2026 at NeurIPS.cc/2026/Workshop/BeNTo · bento-neurips@googlegroups.com