DAEDALUS
Bootstrapping Agent Memory from Self-Generated Tasks
contents
00tl;dr
- LLM agents entering a new environment lack its operational knowledge, and repeat the same mistakes from task to task.
- DAEDALUS builds them a reusable memory with no training tasks and no oracle verifier: an explorer invents tasks that are hard but solvable, a solver attempts them, and each failure yields a heuristic, kept only once the solver succeeds with it repeatedly.
- On AppWorld, τ²-bench and AutomationBench, the memory adds up to +15.9 points of mean success rate over no memory, and raises pass^5 (the share of tasks solved in all five runs) by up to 2.2×.
- That makes it the best method without training tasks, and competitive with methods that learn from curated ones, at a lower inference cost than most of them.
Along with the paper and the code, we also release every trajectory of every agent in every run of the paper at 🤗 illuin/daedalus-traces. Use them as you wish!
01Why agents need memory
LLM agents now carry out real tasks in complex environments, with little human guidance. Doing so reliably takes practical knowledge that carries over from task to task: how the tools behave, how the environment is organised. Without a memory of it, an agent rediscovers its environment on every task, repeats the same mistakes, and needs longer, costlier runs.
Memory methods fix this by learning from past attempts, without touching the model's weights. But they learn from a curated set of training tasks, usually with a verifier that scores each attempt. Building that supervision takes prior knowledge of the environment and a lot of human work, so most environments don't have it, and ever more of them are generated automatically. The alternative, learning from test queries as they arrive, leaves the first tasks with no memory at all.
DAEDALUS removes both requirements. The agent practises before deployment, on tasks it generates itself, and keeps only the lessons that demonstrably help it.
Memory helps, but existing methods need training tasks, and a verifier to grade them, written by people who know the environment. DAEDALUS needs neither.
02How DAEDALUS builds its memory
Before deployment, DAEDALUS spends a fixed budget of practice sessions in the environment. In each one, an Explorer proposes a task with explicit success conditions, and solves it itself to prove it feasible. A Solver, the same model as the agent, then attempts it from a fresh state while an LLM Judge checks each trajectory against the conditions. After every failure, an Extractor writes a heuristic, or revises the last one, and the Solver retries with it in context.
A heuristic is kept only once the Solver succeeds with it three times in a row. A task the Solver never fails is too easy; one it keeps failing is too hard: either way, the Explorer refines it, up to five times, which keeps tasks at the edge of what the Solver can do. A Surveyor spreads the tasks over the environment beforehand, lessons from each refinement become guidelines for the Explorer, and after the last session a Consolidator merges the accepted heuristics into the memory the agent receives at test time.
Apart from the Solver, all these agents are auxiliary agents: they build the memory but are never deployed, so they can run a larger model than the agent.
Here is one real session of our AppWorld run, step by step.
Before the first session, the Surveyor explores AppWorld and sets a coverage goal: a share of tasks per area, Spotify 22%, Todoist 18%, Splitwise 17%, Venmo 16%, Phone 12%… Every session sees the goal and a tally of the tasks accepted so far.
Session 46 of 90. Guided by the coverage goal and its guideline memory, the Explorer spends
19 turns in the apps, then proposes a task, solves it itself to prove it feasible, and
writes the conditions a correct answer must meet.
task T1Please clean up my Todoist by marking done any subtask that's still open even though its parent task is already completed.
success conditionsEvery Todoist subtask that started unfinished under a parent task that was already completed is now completed. No completed Todoist task still has an unfinished subtask.
The Solver attempts T1 from a fresh state each time, and the Judge checks every trajectory against the conditions: ✓✓✓ The Solver succeeds three times in a row, in 10, 14 and 10 turns, and never fails: T1 is too easy, and nothing from it is kept.
The verdict goes back to the Explorer, which makes the task harder by adding the opposite
mismatch, and checks that it is still solvable.
task T2…so task and subtask completion statuses line up: mark done any open subtask whose parent task is already completed, and also mark done any open parent task whose subtasks are all already finished. Tell me how many items you fixed.
success conditionsEvery open subtask under a completed parent, and every open parent whose subtasks were all completed, is now completed. The user is told how many items changed.
The Solver tries the new task T2, from a fresh state and with no heuristic yet. ✓✓✓ It succeeds three times in a row again, in 9, 13 and 8 turns: a mirror case and a count do not make the sweep harder, and T2 is still too easy.
The verdict goes back to the Explorer once more. Rather than add a case, it changes one
rule: a completed task with an unfinished subtask must now be reopened, so the two kinds
of mismatch need opposite fixes.
task T3…so parent tasks and subtasks agree on what's finished: if a task still has an unfinished subtask it shouldn't be completed, and if all its subtasks are already done the task should be completed too. Tell me how many tasks you fixed.
success conditionsEvery task that started out completed with an unfinished subtask is now open, and every task that started out open with all its subtasks done is now completed. The user is told how many parent tasks changed.
The Solver tries T3, from a fresh state and with no heuristic yet, in 9 turns. ✗ The Judge rejects the answer: the Solver checked the parent tasks of a single project, so the mismatches elsewhere in Todoist were never fixed.
The Extractor reads the failed trajectory, without seeing the success conditions, and
writes a heuristic, h1, which the Solver gets in context on its next attempt:
+to reconcile parents with subtasks, inspect every parent with num_sub_tasks > 0, reading its subtasks with todoist.show_sub_tasks before updating anything+a completed parent with an unfinished subtask: reopen it with todoist.update_task(…, is_completed=False)+before completing an open parent, check that all of its subtasks are done+re-read the changed parents and their subtasks, then submit the count with complete_task
✓✓✓ With h1 in context, the Solver succeeds three times in a row, in 8, 23 and 14 turns. A single success could be luck; three consecutive ones show that h1 fixes the failure.
h1 turned a failure into repeated success, so it enters the heuristic memory. Heuristics from tasks that stay too easy or too hard never enter it: this is the only way in.
The session needed refinements, so its history is distilled into the Explorer's
guidelines, which are rewritten. Among them, for every later session:if a cleanup sweep
is still too easy, widen it to one natural consistency invariant across the whole slice,
especially parent/child or group/member agreement […]; don't rely on bolted-on mirror cases
or extra reporting to raise difficulty
After the last of the 90 sessions, 81 heuristics have been accepted. The Consolidator merges them in one call into 65 non-redundant ones, and that memory is injected, whole, at the start of every test task.
03The memory at work
After the last session, the Consolidator merges the accepted heuristics into one memory, and the agent receives it whole at the start of every test task. Here is what that changes on one AppWorld test task, which the agent without memory never solves.
Songs of which genre have I liked the most in my Spotify playlists?
apis.spotify.login(username=…, password=…)apis.spotify.login(username=…, password=…)apis.spotify.show_liked_songs(access_token=…)→ the first page only: 5 of the 25 liked songs, and no playlist in sight✗ off track from hereshow_api_doc(app_name="spotify", api_name="show_playlist_library")apis.spotify.show_song(song_id=…)→ "genre": "classical"liked_in_playlists = [s for s in liked_songs if s["song_id"] in playlist_song_ids]→ 12 songscomplete_task(answer="classical")✗ wrong: the answer is reggaecomplete_task(answer="reggae")✓ rightWhy it works. The task asks for liked songs that are also in the user's playlists: two collections to intersect, and nothing says where either one lives. Without memory, the agent counts the first collection it finds. The memory, learned on generated Spotify tasks and never on this one, says where both live, which field joins them, and to read every page.
04Results
One task shows the mechanism; the benchmarks show how far it carries. We compare DAEDALUS with a no-memory baseline and six memory methods on three benchmarks: app automation in code (AppWorld), customer service conversations (τ²-bench, retail) and SaaS business workflows (AutomationBench, Operations). Five of the methods learn from the benchmark's training tasks, which DAEDALUS never sees; only PREPING, like DAEDALUS, generates its own. The agent is GPT-5.4-mini, with GPT-5.4 for the auxiliary agents (GPT-5.6 Luna and Terra on AutomationBench).
| method | MSR↑ | pass^5↑ | $↓ | MSR↑ | pass^5↑ | $↓ | MSR↑ | pass^5↑ | $↓ |
|---|---|---|---|---|---|---|---|---|---|
| No-memory baseline | 44.3±1.1 | 14.9±2.8 | 3.2 | 57.5±1.6 | 22.5±6.6 | 0.5 | 31.7±3.2 | 12.9±4.0 | 1.0 |
| using training tasks | |||||||||
| AutoGuide | 44.0±0.5 | 17.9±3.0 | 7.6 | 61.5±2.3 | 27.5±7.1 | 3.2 | 35.4±1.7 | 18.6±4.6 | 4.7 |
| ReasoningBank | 48.0±1.0 | 19.6±3.1 | 3.3 | 64.0±2.9 | 32.5±7.5 | 0.5 | 19.1±1.7 | 2.9±2.0 | 0.7 |
| ERL | 58.6±1.3 | 29.2±3.5 | 23.0 | 64.0±3.6 | 30.0±7.2 | 5.1 | 33.4±1.5 | 15.7±4.3 | 3.2 |
| ExpeL | 59.0±1.9 | 33.3±3.6 | 4.6 | 67.0±1.8 | 37.5±7.8 | 0.9 | 42.6±0.5 | 21.4±4.9 | 1.5 |
| ACE | 60.5±0.8 | 41.1±3.8 | 7.6 | 70.5±3.8 | 37.5±7.8 | 2.2 | 26.3±1.9 | 10.0±3.6 | 0.8 |
| DAEDALUS-curated | 60.8±0.9 | 36.3±3.7 | 3.0 | 70.0±3.6 | 35.0±7.5 | 0.6 | 40.0±2.3 | 24.3±5.1 | 1.0 |
| without training tasks | |||||||||
| PREPING | 56.0±1.4 | 25.6±3.4 | 4.0 | 60.0±1.8 | 25.0±6.9 | 0.7 | 34.9±1.5 | 15.7±4.3 | 1.1 |
| DAEDALUS | 60.2±0.9 | 32.1±3.6 | 3.2 | 67.5±2.5 | 37.5±7.8 | 0.5 | 36.0±1.5 | 21.4±4.9 | 1.2 |
- Best without training tasks. +15.9, +10.0 and +4.3 MSR over no memory, pass^5 up to 2.2×, and ahead of PREPING everywhere.
- Competitive with training tasks. Within error of the best such method in four of the six MSR and pass^5 columns.
- No extra inference cost on AppWorld and τ²-bench: the agent needs fewer turns.
- Self-generated tasks recover most of the benefit of curated ones. Run on the training tasks, the same Solver loop (DAEDALUS-curated) ranks first; DAEDALUS is only 0.6, 2.5 and 4.0 MSR points behind it.
05Across model families
So far, the memory has been generated by the agent's own model family. Is it tied to the models that generated it? We also run DAEDALUS with Qwen and DeepSeek model pairs, and give every memory to every base agent.
| memory generated by auxiliary / solver | GPT-5.4-mini | Qwen3.6-35B-A3B | DeepSeek-V4-Flash |
|---|---|---|---|
| no memory (MSR) | 44.3±1.1 | 44.8±1.3 | 80.6±1.6 |
| GPT-5.4 GPT-5.4-mini | +15.9±1.4 | +16.1±1.6 | +8.3±1.8 |
| Qwen3.8-Flash Qwen3.6-35B-A3B | +16.8±2.4 | +29.3±1.7 | +4.9±2.0 |
| DeepSeek-V4-Pro DeepSeek-V4-Flash | +11.4±3.6 | +14.5±1.3 | +3.0±3.0 |
Hover a cell. ◆ a family using its own memory.
All nine gains are positive. GPT-5.4-mini gains as much from the Qwen memory as from its own (+16.8 vs. +15.9); DeepSeek-V4-Flash gains the least, from an 80.6% baseline that leaves little room. With the GPT-5.4 memory, the seven base agents of fig. 1 all succeed more often, in fewer turns.
06What makes generation work
Which parts of the pipeline make the memory work? A cumulative ablation on AppWorld answers it: we rebuild the pipeline one component at a time, each row adding its component to the previous one.
(A) Single Explorer
Heuristics drawn from one 100-turn Explorer trajectory. Exploring alone hurts: the agent does worse than with no memory at all.
(B) Multi-session Explorer
90 sessions of 40 turns, each seeing the heuristics accepted so far. Eighteen times the cost of (A), still below the baseline.
(C) Explorer tasks and Solver traces
The Explorer now proposes tasks, and a heuristic is drawn from every Solver trajectory. +15.8 MSR: what is worth remembering lies in the agent's own attempts.
(D) Solver loop
Heuristics are revised after each failure and kept only after three successes in a row. +4.5 MSR: validation removes the noisy lessons.86.7% of sessions accepted · 183 refinements
(E) Explorer guideline memory
Refinement lessons carried into later sessions: fewer refinements (183 → 161) and 12% cheaper, with MSR within error.91.1% of sessions accepted · 161 refinements
(F) Environment survey (DAEDALUS)
A target distribution of tasks over the environment: 97.8% of sessions end with an accepted heuristic, and generation costs 48% less than (D).97.8% of sessions accepted · 118 refinements
| MSR↑ | pass^5↑ | gen. $↓ | ||
|---|---|---|---|---|
| no memory | 44.3 | █████████ | 14.9 | – |
| (A) Single Explorer | 35.7 | ███████ | 10.1 | $5.2 |
| (B) + Multi-session Explorer | 38.7 | ████████ | 11.3 | $92.0 |
| (C) + Explorer tasks and Solver traces | 54.5 | ███████████ | 25.0 | $46.6 |
| (D) + Solver loop | 59.0 | ████████████ | 30.4 | $210.4 |
| (E) + Explorer guideline memory | 57.9 | ████████████ | 31.5 | $185.3 |
| (F) + Environment survey (DAEDALUS) | 60.2 | ████████████ | 32.1 | $109.7 |
Heuristics must come from the Solver's own attempts: exploring alone makes the agent worse (A, B), Solver traces bring the first jump (C) and validating them the second (D). The guidelines and the survey then make generation cheaper rather than better. With fewer refinements, the Solver and the Judge cost 60% less, and the whole run costs half as much as (D).
- explorer
- solver
- judge
- extractor
07Using the memory
Generating good heuristics is half of the story; the other half is how the agent receives them. Heuristics accepted in different sessions overlap. Consolidated into one memory and injected whole at task start, they are both the best and the cheapest option, 10.7 MSR points above the raw heuristics. Retrieving five heuristics per turn helps less, and BM25, embeddings and a random pick are all within error of each other. Retrieval ranks the heuristics by their resemblance to the agent's current turn, but the heuristic the agent needs often reads nothing like that turn: it warns about a mistake still to come, not about what the agent is doing now.
Retrieving more per turn helps, but never reaches the whole memory at start.
- mean success rate
- pass^5
- whole memory at start
08Ranking models with generated tasks
The generated tasks have a use of their own, beyond memory. Environments without training tasks usually lack a test set as well, which makes it hard to compare agents on them. The tasks DAEDALUS generates could fill that role.
To test it, we evaluate nine models, without memory, on two task sets: the 81 tasks accepted during a 90-session generation run on AppWorld, and AppWorld's official test_normal split. If the generated tasks are a good proxy, the models should come out in the same order on both.
- nemotron-3.5-lightning
- gpt-oss-120b
- minimax-m2.7
- gpt-5.4-mini
- qwen3.6-35b-a3b
- gemma-4-31b-it
- mimo-v2.5
- deepseek-v4-flash
- gpt-5.6-luna
data table
model MSR aw MSR gen pass³ aw pass³ gen ──────────────────────────────────────────────────────────────────────── nemotron-3.5-lightning 6.5 13.6 0.2 1.2 gpt-oss-120b 19.2 39.9 3.2 12.3 minimax-m2.7 25.1 43.6 10.9 21.0 gpt-5.4-mini 44.3 66.7 22.2 40.7 qwen3.6-35b-a3b 44.8 67.9 27.7 46.9 gemma-4-31b-it 51.0 61.3 32.4 33.3 mimo-v2.5 52.6 76.0 33.2 51.2 deepseek-v4-flash 60.8 81.5 37.4 61.7 gpt-5.6-luna 66.2 90.1 49.7 82.7 aw = AppWorld test_normal · gen = DAEDALUS-generated tasks · values in %
The rankings agree. Kendall's τ is 0.89 for mean success rate and 0.89 for pass³ (both p < .001), and 34 of the 36 pairs of models are ordered the same way on both sets.
Absolute scores differ: every model does better on the generated tasks, which were calibrated to GPT-5.4-mini, the solver that generated them. But the order of models is preserved, so tasks generated by DAEDALUS can serve as a proxy for ranking models when no curated test set exists.
09Key takeaways
- No training tasks, no verifier. DAEDALUS builds an agent's memory from tasks it generates itself.
- Up to +15.9 points of mean success rate over no memory: the strongest method without training tasks, and competitive with those that learn from curated ones, at little or no extra inference cost.
- Validation is what works. Heuristics grounded in the agent's own failures and confirmed by its successes help; heuristics from exploration alone hurt.
- The memory transfers across model families.
- Give the whole memory at task start. Consolidated and injected once, it beats per-turn retrieval.
- The generated tasks rank models in the same order as AppWorld's official test set (τ = 0.89).
@Citation
If DAEDALUS, its code or its artefacts are useful to you, please cite:
@misc{edy2026daedalus,
title = {{DAEDALUS}: Bootstrapping Agent Memory from Self-Generated Tasks},
author = {Edy, Antoine and Conti, Max and Xing, Victor and Allard, Marc-Antoine and Benhamdane, Nawfal and Viaud, Gautier},
year = {2026},
eprint = {2610.08048},
archivePrefix = {arXiv},
url = {https://arxiv.org/abs/2610.08048}
}AAppendix
three more analyses: the judge, the budget, the auxiliary model
Can the judge be trusted?
The loop rests on an LLM judge. It agrees substantially with each benchmark's own verifier (κ > 0.7), and five calls on the same trajectory agree with each other (κ > 0.8). Its precision beats its recall everywhere: it rejects some successes rather than accepting failures, the safe side for accepting heuristics.
| N | precision | recall | κ bench | κ inter | |
|---|---|---|---|---|---|
| AppWorld | 168 | .943 | .846 | .807 | .848 |
| τ²-bench | 40 | .889 | .842 | .749 | 1.000 |
| AutomationBench | 70 | .863 | .807 | .729 | .902 |
How many sessions?
Five sessions already bring more than half of the gain of the 90-session peak, so a useful memory is cheap. Past 90 sessions, success drops while the memory keeps growing.
- mean success rate
- pass^5
- heuristics in memory
Does it need a large model?
With GPT-5.4-mini for every auxiliary agent, the solver's own model, DAEDALUS still adds 6.9 MSR points: about half of the gain with GPT-5.4.
| auxiliary model | MSR↑ | pass^5↑ | gen. $↓ |
|---|---|---|---|
| no memory | 44.3±1.1 | 14.9±2.8 | – |
| GPT-5.4-mini | 51.2±1.6 | 20.2±3.1 | $67.2 |
| GPT-5.4 | 60.2±0.9 | 32.1±3.6 | $109.7 |