Agent System Trilemma
- performance (, task success and accuracy)
- efficiency (, time or steps required to complete tasks)
- cost (., computational and monetary resources consumed)
EvoRoute
A self-evolving routing paradigm that dismantles the trilemma through fine-grained model selection. Before executing each step, it dynamically selects the most judicious LLM by:
- retrieval: performing a multifaceted retrieval to identify historically analogous sub-task executions from an evolving knowledge base;
- filtration: distilling a Pareto-optimal set of candidate models, , those that are not dominated across the axes of cost, efficiency, and performance;
- selection: leveraging a lightweight decision model to make the final selection based on this rich, context-aware statistical evidence.
Notations
- : the designed complex agentic AI system
- : the set of agent roles (, web-browser, coder)
- : the pool of available LLM backbones
- : system state, typically implemented as a shared memory or scratchpad
- : a set of external tools, such as code interpreters or web search APIs
- : the full action space, including both natural language actions and tool invocations, formally
- governs the transition dynamics of the system
- selects the active agent at each time step
Objective Formulation
- : the dynamic routing policy that selects an LLM for the active agent at each step
- : task performance
- : the cumulative monetary and computational expenditure
- : the total wall-clock execution time
- : the full execution trajectory of the system
Methodology
Step-level experience base
The backbone of EvoRoute is an evolving knowledge base built from prior executions. After a task finishes, the full trajectory is split into step-level records:
Each record stores:
- : active agent role,
- : LLM used at this step,
- : sub-task instruction,
- : embedding of the instruction,
- : tools used,
- : cost,
- : wall-clock duration,
- : whether the step executed successfully,
- : final task-level success signal.
After each run:
Multi-Faceted Retrieval
When a new step arrives, EvoRoute retrieves relevant historical records from . Instead of relying on one notion of similarity, it uses three.
-
Agent Role Match
-
Semantic Similarity Retrieval
- is implemented via MiniLM
-
Tool Congruence Retrieval
uses a two-stage predictor:
- Keyword heuristic
- Using a predefined dictionary to map explicit trigger keywords(e.g., "search" for ; "run", "plot" for )
- Cheap LLM fallback
- if heuristics fail, use Qwen3-14B in zero-shot mode
- Keyword heuristic
The final candidate set:
Pareto-Optimal Filtration and Selection
From the retrieved records, EvoRoute extracts candidate models:
For each candidate model , it estimates:
- average performance
- average cost
- average delay
A model is dominated if another model exists that is superior or equal on all three axes and strictly superior on at least one.
Retaining only the non-dominated models, we form the Pareto-optimal set,
Thompson-sampling-based model selection
If EvoRoute always picked the current best average, it would become too greedy and stop learning.
It assumes each metric follows a Normal distribution and models the uncertainty over its mean and variance using a Normal-Inverse-Gamma conjugate prior.
First, compute the sample statistics for each metric : the count , the sample mean and the sample variance . These statistics are used to parameterize the NIG posteriors, NIG(), where , , , and
At decision time, it samples a stochastic utility:
and selects:
where reflect the desired trilemma trade-off (, , )
Crucially, this selection is not the end of the process. Once the agent powered by completes its action, the observed outcome is logged back into the knowledge base . This closes the feedback loop, ensuring that every decision and its outcome contribute to the system’s ever-improving wisdom, thereby realizing the self-evolving nature of EvoRoute.