DROID: Dual Representation for Out-of-Scope Intent Detection
Wael Rashwan, Hossam M. Zawbaa, Sourav Dutta, Haytham Assem
TL;DR
DROID tackles out-of-scope (OOS) intent detection in task-oriented dialogue by fusing two complementary sentence encoders—a general-purpose Universal Sentence Encoder (USE) and a domain-adapted TSDAE—within a lightweight end-to-end framework. A single calibrated threshold on in-domain validation differentiates known intents from OOS without post-hoc scoring or strong distributional assumptions, while synthetic feature-space outliers and open-domain negatives strengthen boundary learning. With only 1.56M trainable parameters, DROID achieves state-of-the-art macro-F1 on Known and Unknown intents across CLINC-150, BANKING77, and STACKOVERFLOW, demonstrating robustness in low-resource settings and under varying known-intent coverage. The approach offers an efficient, deployable alternative to large LLM-based baselines, highlighting the benefits of dual representations and explicit confidence calibration for open-world intent understanding in real-time systems.
Abstract
Detecting out-of-scope (OOS) user utterances remains a key challenge in task-oriented dialogue systems and, more broadly, in open-set intent recognition. Existing approaches often depend on strong distributional assumptions or auxiliary calibration modules. We present DROID (Dual Representation for Out-of-Scope Intent Detection), a compact end-to-end framework that combines two complementary encoders -- the Universal Sentence Encoder (USE) for broad semantic generalization and a domain-adapted Transformer-based Denoising Autoencoder (TSDAE) for domain-specific contextual distinctions. Their fused representations are processed by a lightweight branched classifier with a single calibrated threshold that separates in-domain and OOS intents without post-hoc scoring. To enhance boundary learning under limited supervision, DROID incorporates both synthetic and open-domain outlier augmentation. Despite using only 1.5M trainable parameters, DROID consistently outperforms recent state-of-the-art baselines across multiple intent benchmarks, achieving macro-F1 improvements of 6--15% for known and 8--20% for OOS intents, with the most significant gains in low-resource settings. These results demonstrate that dual-encoder representations with simple calibration can yield robust, scalable, and reliable OOS detection for neural dialogue systems.
