A NeurIPS 2026 Workshop · Sydney, Australia

BeNTo: Beyond Next‑Token Prediction —
Diffusion & Flow Models for Next‑Generation Decoding

Beyond next-token prediction — exploring the theory, algorithms, and applications of discrete diffusion and flow models for parallel, non-causal generation.

Submissions dueAug 29, 2026
NotificationSep 28, 2026
Workshop · single-dayDec 11 or 12, 2026

01 Overview

Rethinking the order of generation.

Autoregressive models produce their output one element at a time, in a fixed order. The recipe is remarkably general, and it has carried the field a long way. However, it also ties the order of computation to the order of the result, leaving parallelism, revision, and control hard to reach. Discrete diffusion and flow models take a different route. They generate by iterative denoising: refining many positions at once, and revisiting earlier choices as a sample takes shape.

That shift brings a different set of computational properties — non-causal, parallel, and controllable — along with open questions spanning mathematics, algorithms, and engineering, and a growing body of work that combines the two paradigms rather than choosing between them. This workshop joins two conversations: the theory and algorithms that make discrete generative models work, and the applications and systems that put them to use.

Track 01

Depth — Theories & Algorithms

Discrete diffusion and flow models at the intersection of generative modeling, optimal transport, and stochastic optimal control.

Generative Modeling
Discrete diffusion & flow models, optimal transport, optimal control, Schrödinger bridges.
Probabilistic Inference
Discrete diffusion samplers, adjoint-based samplers, advanced MCMC, variational inference.
Track 02

Breadth — Applications & Systems

Novel applications and scalable systems that exploit non-causal, parallel generation.

Applications
AI for science, multimodal generation, reward alignment, inverse problems, benchmarks & datasets.
Systems & Empirical Analysis
Foundation models, large-scale training/inference, network architectures.

02 Scope

Topics of interest.

This workshop considers, but is not limited to, the following topics.

IDiscrete diffusion & flow models
IIOptimal transport
IIIOptimal & stochastic control
IVSchrödinger bridges
VDiscrete diffusion & adjoint-based samplers
VIMCMC & variational inference
VIIAI for science
VIIIMultimodal generation
IXReward alignment
XInverse problems
XIBenchmarks & datasets
XIIFoundation models & large-scale training/inference
XIIIArchitectures for non-autoregressive generation
XIVSystems & hardware for diffusion models

03 Call for Papers

Share your work.

What to submit

We invite submissions on any topic within the workshop's scope, across both tracks, from theory and algorithms to applications and systems. All submissions are handled through OpenReview.

  • Page limit. Submissions may be either 4 or 8 pages, excluding references and appendices.
  • Peer review. Each submission receives 3 reviews; each reviewer handles at most 3 papers.
  • Mutual-review policy. At least one qualified author per submission serves as a reviewer.
  • Travel grants for students and early-career researchers with accepted papers.

Presentation

Accepted papers are presented as posters, with a selection of contributed orals. A Best Paper Award recognizes outstanding work.

04 Important Dates

Mark your calendar.

Jul 20, 2026Call for Papers & submissions open OpenReview
Aug 29, 2026Paper submission deadline
Aug 30 – Sep 20, 2026Peer review period
Sep 21 – 25, 2026Meta-review & discussion
Sep 28, 2026Accept / reject notification
Oct 20, 2026Camera-ready deadline tentative
Dec 11 or 12, 2026Workshop @ NeurIPS 2026 single-day; exact day TBD

All deadlines are 23:59 AoE. Dates are indicative and subject to change.

05 Schedule

A one-day program.

A full day of invited talks, contributed orals, two poster sessions, and a panel. Speaker slot assignments are being finalized.

08:50–09:00Opening remarksWelcome
09:00–09:30Invited Talk 1Keynote
09:30–10:00Invited Talk 2Keynote
10:00–11:00Poster session 1 & breakPosters
11:00–11:30Invited Talk 3Keynote
11:30–12:00Invited Talk 4Keynote
12:00–12:30Oral session 1 (3 papers)Orals
12:30–13:30LunchBreak
13:30–14:30Panel discussionPanel
14:30–15:00Invited Talk 5Keynote
15:00–15:30Invited Talk 6Keynote
15:30–16:00Oral session 2 (3 papers)Orals
16:00–17:00Poster session 2 & socialPosters
17:00–17:10Awards & closingBest Paper

Tentative program · times shown in the local conference time zone. Invited-talk slots will be mapped to speakers closer to the event.


06 Invited Speakers

Six voices shaping the field.

Talk titles are tentative and may be refined before the workshop.

SEPortrait of Stefano Ermon
Stefano Ermon
Stanford University

Talk TitleProduction-Grade Diffusion LLMs

Associate Professor of Computer Science at Stanford, where his group advances generative modeling, probabilistic inference, and diffusion methods. His recent work pushes diffusion language models toward production-scale text generation.

PMPortrait of Pavlo Molchanov
Pavlo Molchanov
NVIDIA Research · Research Director

Talk TitleEfficient Diffusion Language Models from Autoregressive LLMs

Research Director at NVIDIA Research, focused on efficient deep learning — model compression, acceleration, and foundation models. His talk explores distilling fast diffusion language models from autoregressive LLMs.

RBPortrait of Rianne van den Berg
Rianne van den Berg
Microsoft Research · Sr. Principal Research Manager

Talk TitleDiffusion Models for the Natural Sciences

Senior Principal Research Manager at Microsoft Research, known for foundational work on normalizing flows and discrete diffusion. She applies generative models to problems across the natural sciences.

SCPortrait of Sitan Chen
Sitan Chen
Harvard University

Talk TitleEliciting Non-Causal Reasoning in Diffusion Language Models

Assistant Professor at Harvard working on the theoretical foundations of machine learning, including the analysis of diffusion models and sampling algorithms, and how non-causal generation enables new forms of reasoning.

CLPortrait of Chongxuan Li
Chongxuan Li
Renmin University · Gaoling School of AI

Talk TitleLLaDA: A New Paradigm for Large Language Modeling

Associate Professor at the Gaoling School of AI, Renmin University of China, specializing in deep generative models and diffusion. His group developed LLaDA, a large language diffusion model.

JSPortrait of Jiaxin Shi
Jiaxin Shi
Meta Superintelligence Labs

Talk TitleSynthesizing Autoregression and Diffusion for Flexible Generation and Planning

Researcher at Meta Superintelligence Labs working on probabilistic machine learning and generative models, unifying autoregression and diffusion for flexible, controllable generation and planning.


07 Panel

A conversation on what comes next.

A moderated discussion on the open problems and near-term future of diffusion and flow models.

MSPortrait of Minhyuk Sung
Moderator
Minhyuk Sung
KAIST · Workshop Organizer
GLPortrait of Guan-Horng Liu
Guan-Horng Liu
Meta Superintelligence Labs
AGPortrait of Aditya Grover
Aditya Grover
UCLA
AVPortrait of Arash Vahdat
Arash Vahdat
NVIDIA Research
ADPortrait of Arnaud Doucet
Arnaud Doucet
Google DeepMind

08 Organizers

Brought to you by the organizing team.

General enquiries: bento-neurips@googlegroups.com

MSPortrait of Minhyuk Sung
Minhyuk Sung
KAIST

Associate Professor at KAIST working on 3D vision, computer graphics, and generative models.

JKPortrait of Jaihoon Kim
Jaihoon Kim
KAIST

PhD student at KAIST, working on diffusion models and generative AI.

MTPortrait of Molei Tao
Molei Tao
Georgia Tech

Professor at Georgia Tech working on Generative AI such as diffusion models.

PCPortrait of Pranam Chatterjee
Pranam Chatterjee
University of Pennsylvania

Assistant Professor at the University of Pennsylvania developing generative models for biological and therapeutic design.

STPortrait of Sophia Tang
Sophia Tang
University of Pennsylvania

Undergraduate Researcher at the University of Pennsylvania researching generative models for biomolecular design.

NDPortrait of Nolan Dey
Nolan Dey
Cerebras Systems

Research Scientist at Cerebras Systems working on large-scale training, scaling laws, and efficient deep learning.

SSPortrait of Subham Sekhar Sahoo
Subham Sekhar Sahoo
MBZUAI · US Labs

Researcher at MBZUAI focused on discrete diffusion language models and efficient generative modeling.

09 Sponsors & Support

Support that widens the door.

Sponsorship helps us fund travel grants and broaden participation for students and early-career researchers. We gratefully acknowledge the organizations supporting this workshop.

Ready to submit?

Join us at NeurIPS 2026 to push generation beyond next-token prediction. Submissions open July 20 and close August 29, 2026.

Submit on OpenReview

Submissions open Jul 20, 2026 at NeurIPS.cc/2026/Workshop/BeNTo · bento-neurips@googlegroups.com