Junyuan Zheng, Onkar Salvi, John Chan
Dialogue state tracking is the component that remembers what a conversation has established so far. At every turn it updates its picture of what the user wants, and everything downstream, deciding what the system should do and what it should say, reads from that picture.
The difficulty comes from scale rather than from any single conversation. Assistants like Siri, Alexa and Google Assistant sit in front of a large and growing catalogue of services and APIs, each with its own vocabulary of things a user might specify. A tracker that learned one service's slots has to cope with services it was never trained on.
Attending over candidates is a sensible response to that. Rather than treating the state as a fixed set of fields to fill, the model considers the candidate values actually available in context, which is what allows a new service to be handled without a new model.
Dialogue state tracking (DST) is one of the core components in task-oriented dialogue systems. At each turn in a conversation, DST estimates the user belief or dialogue state, which is used as input for downstream modules to predict system actions and generate responses. The increasingly popular dialogue system applications like Google Assistant, Siri and Alexa need to support a large number of services and APIs, resulting in growing attention to the scalability of such…
DPNeXt: A Lightweight Multi-Scale Feature Fusion Framework for Efficient ViT-Based Multi-Task Dense Prediction
arXiv (cs.CL) · July 17, 2026AI Watermark Evidence Fails Forensic Readiness: An Empirical Evaluation
arXiv (cs.CV) · July 17, 2026Beyond Unfolding: 60x Faster One-Stage Unmixing for Closely-Spaced Infrared Small Targets
arXiv (cs.AI) · July 17, 2026Robustness of Reinforcement Learning-Based Congestion Management in Low-Voltage Grids