FlashResearch: Real-time Agent Orchestration for Efficient Deep Research

Lunyiu Nie; Nedim Lipka; Ryan A. Rossi; Swarat Chaudhuri

FlashResearch: Real-time Agent Orchestration for Efficient Deep Research

Lunyiu Nie, Nedim Lipka, Ryan A. Rossi, Swarat Chaudhuri

TL;DR

FlashResearch tackles latency and adaptability bottlenecks in deep research by converting sequential reasoning into a parallel, tree-structured workflow. It introduces an adaptive planner, a real-time orchestrator, and a multi-dimensional parallelization engine to dynamically expand and prune a research tree while speculative execution proceeds across breadth and depth under a time budget $t_{\text{max}}$. Empirical results on DeepResearchGym and DeepResearch Bench show up to a 5× speed-up with comparable or improved quality, validating the approach under interactive budgets. The framework advances practical, responsive deep research by enabling real-time replanning, resource reallocation, and asynchronous, parallel exploration of multiple research avenues.

Abstract

Deep research agents, which synthesize information across diverse sources, are significantly constrained by their sequential reasoning processes. This architectural bottleneck results in high latency, poor runtime adaptability, and inefficient resource allocation, making them impractical for interactive applications. To overcome this, we introduce FlashResearch, a novel framework for efficient deep research that transforms sequential processing into parallel, runtime orchestration by dynamically decomposing complex queries into tree-structured sub-tasks. Our core contributions are threefold: (1) an adaptive planner that dynamically allocates computational resources by determining research breadth and depth based on query complexity; (2) a real-time orchestration layer that monitors research progress and prunes redundant paths to reallocate resources and optimize efficiency; and (3) a multi-dimensional parallelization framework that enables concurrency across both research breadth and depth. Experiments show that FlashResearch consistently improves final report quality within fixed time budgets, and can deliver up to a 5x speedup while maintaining comparable quality.

FlashResearch: Real-time Agent Orchestration for Efficient Deep Research

TL;DR

Abstract

FlashResearch: Real-time Agent Orchestration for Efficient Deep Research

TL;DR

Abstract

Paper Structure

Table of Contents

Figures (3)