DAEDALUS

Bootstrapping Agent Memory from Self-Generated Tasks

Antoine Edy*† · Max Conti* · Victor Xing* · Marc-Antoine Allard · Nawfal Benhamdane · Gautier Viaud

Illuin Technology
* equal contribution · † corresponding author
contents
  1. 00tl;dr
  2. 01Why agents need memory
  3. 02How DAEDALUS builds its memory
  4. 03The memory at work
  5. 04Results
  6. 05Across model families
  7. 06What makes generation work
  8. 07Using the memory
  9. 08Ranking models with generated tasks
  10. 09Key takeaways
  11. @Citation
  12. AAppendix
myth the names behind DAEDALUS

Why DAEDALUS?

In the myth, Daedalus builds the Labyrinth of Crete for King Minos, and he alone knows that it has a way out. DAEDALUS casts the same characters.

  • sourceDaedalus──▶ ExplorerThe architect of the Labyrinth. He built it so intricate that, Ovid says, he could barely find his own way back out; but he could. The Explorer builds each task hard, and solves it first to prove there is a way out.
  • sourceTheseus──▶ SolverWalks into the Labyrinth through its entrance, with no map. The Solver attempts the task from a fresh start, knowing only what the memory tells it.
  • sourcethe Minotaur──▶ the errorWhat waits inside and ends the attempt. Every lesson in the memory starts from one of these failures.
  • sourceMinos──▶ JudgeThe king of Crete, and after his death a judge of the dead. The Judge decides whether an attempt met the task's conditions.
  • sourceAriadne──▶ ExtractorOn Daedalus's advice, she hands Theseus a thread to find his way back. After each failure, the Extractor writes the heuristic the Solver takes into its next attempt.
  • sourcethe thread──▶ a heuristic“Clue” comes from “clew”, a ball of thread, after this story. A heuristic is kept only once it has led the Solver out three times in a row.
  • sourceIcarus──▶ refinementDaedalus told his son to fly neither too low, where the sea soaks the wings, nor too high, where the sun melts the wax. A task the Solver never fails flies too low, one it always fails too high: the Explorer refines it back to the middle.
  • sourceTalos──▶ SurveyorThe bronze guardian who went round Crete three times a day. The Surveyor goes round the whole environment first, so that the tasks cover all of it.
  • the clew──▶ ConsolidatorThe Consolidator winds the kept threads into one ball: the memory the agent carries into every test task.

Didn't get it? Read the blog post first, then come back!

fig. 1 Using DAEDALUS on 7 base agents (AppWorld)
One DAEDALUS memory, generated on AppWorld, makes seven base agents from six model families succeed more often, in fewer turns. AppWorld test_normal, five runs each. Hollow: no memory; filled: with the DAEDALUS memory. Up is more success, left is fewer turns.

00tl;dr

  • LLM agents entering a new environment lack its operational knowledge, and repeat the same mistakes from task to task.
  • DAEDALUS builds them a reusable memory with no training tasks and no oracle verifier: an explorer invents tasks that are hard but solvable, a solver attempts them, and each failure yields a heuristic, kept only once the solver succeeds with it repeatedly.
  • On AppWorld, τ²-bench and AutomationBench, the memory adds up to +15.9 points of mean success rate over no memory, and raises pass^5 (the share of tasks solved in all five runs) by up to 2.2×.
  • That makes it the best method without training tasks, and competitive with methods that learn from curated ones, at a lower inference cost than most of them.

Along with the paper and the code, we also release every trajectory of every agent in every run of the paper at 🤗 illuin/daedalus-traces. Use them as you wish!

01Why agents need memory

LLM agents now carry out real tasks in complex environments, with little human guidance. Doing so reliably takes practical knowledge that carries over from task to task: how the tools behave, how the environment is organised. Without a memory of it, an agent rediscovers its environment on every task, repeats the same mistakes, and needs longer, costlier runs.

Memory methods fix this by learning from past attempts, without touching the model's weights. But they learn from a curated set of training tasks, usually with a verifier that scores each attempt. Building that supervision takes prior knowledge of the environment and a lot of human work, so most environments don't have it, and ever more of them are generated automatically. The alternative, learning from test queries as they arrive, leaves the first tasks with no memory at all.

DAEDALUS removes both requirements. The agent practises before deployment, on tasks it generates itself, and keeps only the lessons that demonstrably help it.

fig. 2 three ways to meet a new environment
before deploymenttest task 1test task 2
no memory · every task starts from zero
nothing: the agent meets its environment on the first test task
agent ──▶ error A✗ failed
agent ──▶ error A✗ failedthe same mistake, again
memory methods · learn from curated training tasks
train 1 ›✗ error A → lesson Atrain 2 ›✗ error B → lesson Bneeds training tasks + a verifier
agent ──▶ error A✓ solved
agent ──▶ error B✓ solved
DAEDALUS · learns from tasks it writes itself
no training tasks, no human priors
agent ──▶ error A✓ solved
agent ──▶ error B✓ solved

Memory helps, but existing methods need training tasks, and a verifier to grade them, written by people who know the environment. DAEDALUS needs neither.

02How DAEDALUS builds its memory

Before deployment, DAEDALUS spends a fixed budget of practice sessions in the environment. In each one, an Explorer proposes a task with explicit success conditions, and solves it itself to prove it feasible. A Solver, the same model as the agent, then attempts it from a fresh state while an LLM Judge checks each trajectory against the conditions. After every failure, an Extractor writes a heuristic, or revises the last one, and the Solver retries with it in context.

A heuristic is kept only once the Solver succeeds with it three times in a row. A task the Solver never fails is too easy; one it keeps failing is too hard: either way, the Explorer refines it, up to five times, which keeps tasks at the edge of what the Solver can do. A Surveyor spreads the tasks over the environment beforehand, lessons from each refinement become guidelines for the Explorer, and after the last session a Consolidator merges the accepted heuristics into the memory the agent receives at test time.

Apart from the Solver, all these agents are auxiliary agents: they build the memory but are never deployed, so they can run a larger model than the agent.

Here is one real session of our AppWorld run, step by step.

fig. 3 one generation session, step by step
01 surveyor · maps the environment once

Before the first session, the Surveyor explores AppWorld and sets a coverage goal: a share of tasks per area, Spotify 22%, Todoist 18%, Splitwise 17%, Venmo 16%, Phone 12%… Every session sees the goal and a tally of the tasks accepted so far.

02 explorer · proposes a task, T1

Session 46 of 90. Guided by the coverage goal and its guideline memory, the Explorer spends 19 turns in the apps, then proposes a task, solves it itself to prove it feasible, and writes the conditions a correct answer must meet. task T1Please clean up my Todoist by marking done any subtask that's still open even though its parent task is already completed. success conditionsEvery Todoist subtask that started unfinished under a parent task that was already completed is now completed. No completed Todoist task still has an unfinished subtask.

03 solver+judge · T1 is too easy

The Solver attempts T1 from a fresh state each time, and the Judge checks every trajectory against the conditions: ✓✓✓ The Solver succeeds three times in a row, in 10, 14 and 10 turns, and never fails: T1 is too easy, and nothing from it is kept.

04 explorer · refinement #1, T2

The verdict goes back to the Explorer, which makes the task harder by adding the opposite mismatch, and checks that it is still solvable. task T2…so task and subtask completion statuses line up: mark done any open subtask whose parent task is already completed, and also mark done any open parent task whose subtasks are all already finished. Tell me how many items you fixed. success conditionsEvery open subtask under a completed parent, and every open parent whose subtasks were all completed, is now completed. The user is told how many items changed.

05 solver+judge · T2 is too easy

The Solver tries the new task T2, from a fresh state and with no heuristic yet. ✓✓✓ It succeeds three times in a row again, in 9, 13 and 8 turns: a mirror case and a count do not make the sweep harder, and T2 is still too easy.

06 explorer · refinement #2, T3

The verdict goes back to the Explorer once more. Rather than add a case, it changes one rule: a completed task with an unfinished subtask must now be reopened, so the two kinds of mismatch need opposite fixes. task T3…so parent tasks and subtasks agree on what's finished: if a task still has an unfinished subtask it shouldn't be completed, and if all its subtasks are already done the task should be completed too. Tell me how many tasks you fixed. success conditionsEvery task that started out completed with an unfinished subtask is now open, and every task that started out open with all its subtasks done is now completed. The user is told how many parent tasks changed.

07 solver+judge · T3 fails

The Solver tries T3, from a fresh state and with no heuristic yet, in 9 turns. ✗ The Judge rejects the answer: the Solver checked the parent tasks of a single project, so the mismatches elsewhere in Todoist were never fixed.

08 extractor · writes h1

The Extractor reads the failed trajectory, without seeing the success conditions, and writes a heuristic, h1, which the Solver gets in context on its next attempt: +to reconcile parents with subtasks, inspect every parent with num_sub_tasks > 0, reading its subtasks with todoist.show_sub_tasks before updating anything+a completed parent with an unfinished subtask: reopen it with todoist.update_task(…, is_completed=False)+before completing an open parent, check that all of its subtasks are done+re-read the changed parents and their subtasks, then submit the count with complete_task

09 solver+judge · three successes in a row

✓✓✓ With h1 in context, the Solver succeeds three times in a row, in 8, 23 and 14 turns. A single success could be luck; three consecutive ones show that h1 fixes the failure.

10 memory · h1 is accepted

h1 turned a failure into repeated success, so it enters the heuristic memory. Heuristics from tasks that stay too easy or too hard never enter it: this is the only way in.

11 guidelines · a lesson for the Explorer

The session needed refinements, so its history is distilled into the Explorer's guidelines, which are rewritten. Among them, for every later session:if a cleanup sweep is still too easy, widen it to one natural consistency invariant across the whole slice, especially parent/child or group/member agreement […]; don't rely on bolted-on mirror cases or extra reporting to raise difficulty

12 consolidator · one memory for test time

After the last of the 90 sessions, 81 heuristics have been accepted. The Consolidator merges them in one call into 65 non-redundant ones, and that memory is injected, whole, at the start of every test task.

01/12or ← → keys

03The memory at work

After the last session, the Consolidator merges the accepted heuristics into one memory, and the agent receives it whole at the start of every test task. Here is what that changes on one AppWorld test task, which the agent without memory never solves.

fig. 4 one test task, turn by turn, without and with memoryAppWorld test taskSongs of which genre have I liked the most in my Spotify playlists?
without memorywith memory
t1–2
Lists the Spotify APIs and logs in.apis.spotify.login(username=…, password=…)
t1–2
Lists the Spotify APIs and logs in.apis.spotify.login(username=…, password=…)
t3
Fetches the liked songs, without paging arguments.apis.spotify.show_liked_songs(access_token=…)→ the first page only: 5 of the 25 liked songs, and no playlist in sight✗ off track from here
t3
memory › Your playlists come from show_playlist_library, which lists their song_ids; liked songs carry a song_id. Compare by ID.Reads the docs of the two collections the heuristic names.show_api_doc(app_name="spotify", api_name="show_playlist_library")
t4–5
Looks for the genre field: one failed call, then finds it on each song.apis.spotify.show_song(song_id=…)→ "genre": "classical"
t4
memory › Fully paginate Spotify collections (page_limit ≤ 20) before comparing sets.Pages through the 8 playlists (57 song IDs) and all 25 liked songs, and keeps the liked songs that are in a playlist.liked_in_playlists = [s for s in liked_songs if s["song_id"] in playlist_song_ids]→ 12 songs
t6
Counts genres over its 5 liked songs, never opening a playlist.→ classical 2 · EDM 1 · rock 1 · jazz 1
t4
Counts genres over exactly those 12 songs.→ reggae 4 · EDM 2 · classical 2 · rock 1 · pop 1 · …
t7
Answers from the wrong set of songs.complete_task(answer="classical")✗ wrong: the answer is reggae
t5
Answers, two turns earlier.complete_task(answer="reggae")✓ right
×5
✗ 0/5 runsanswers R&B, EDM, classical; median 9 turns
×5
✓ 5/5 runsreggae every time; median 7 turns

Why it works. The task asks for liked songs that are also in the user's playlists: two collections to intersect, and nothing says where either one lives. Without memory, the agent counts the first collection it finds. The memory, learned on generated Spotify tasks and never on this one, says where both live, which field joins them, and to read every page.

04Results

One task shows the mechanism; the benchmarks show how far it carries. We compare DAEDALUS with a no-memory baseline and six memory methods on three benchmarks: app automation in code (AppWorld), customer service conversations (τ²-bench, retail) and SaaS business workflows (AutomationBench, Operations). Five of the methods learn from the benchmark's training tasks, which DAEDALUS never sees; only PREPING, like DAEDALUS, generates its own. The agent is GPT-5.4-mini, with GPT-5.4 for the auxiliary agents (GPT-5.6 Luna and Terra on AutomationBench).

table 1 main results, five inference runs per method
methodMSR↑pass^5↑$↓MSR↑pass^5↑$↓MSR↑pass^5↑$↓
No-memory baseline44.3±1.114.9±2.83.257.5±1.622.5±6.60.531.7±3.212.9±4.01.0
using training tasks
AutoGuide44.0±0.517.9±3.07.661.5±2.327.5±7.13.235.4±1.718.6±4.64.7
ReasoningBank48.0±1.019.6±3.13.364.0±2.932.5±7.50.519.1±1.72.9±2.00.7
ERL58.6±1.329.2±3.523.064.0±3.630.0±7.25.133.4±1.515.7±4.33.2
ExpeL59.0±1.933.3±3.64.667.0±1.837.5±7.80.942.6±0.521.4±4.91.5
ACE60.5±0.841.1±3.87.670.5±3.837.5±7.82.226.3±1.910.0±3.60.8
DAEDALUS-curated60.8±0.936.3±3.73.070.0±3.635.0±7.50.640.0±2.324.3±5.11.0
without training tasks
PREPING56.0±1.425.6±3.44.060.0±1.825.0±6.90.734.9±1.515.7±4.31.1
DAEDALUS60.2±0.932.1±3.63.267.5±2.537.5±7.80.536.0±1.521.4±4.91.2
MSR: mean success rate (%); pass^5: tasks solved in all five runs (%); $: inference cost per run (USD), memory generation excluded. Bold: best in the column; underlined: best without training tasks.
  • Best without training tasks. +15.9, +10.0 and +4.3 MSR over no memory, pass^5 up to 2.2×, and ahead of PREPING everywhere.
  • Competitive with training tasks. Within error of the best such method in four of the six MSR and pass^5 columns.
  • No extra inference cost on AppWorld and τ²-bench: the agent needs fewer turns.
  • Self-generated tasks recover most of the benefit of curated ones. Run on the training tasks, the same Solver loop (DAEDALUS-curated) ranks first; DAEDALUS is only 0.6, 2.5 and 4.0 MSR points behind it.

05Across model families

So far, the memory has been generated by the agent's own model family. Is it tied to the models that generated it? We also run DAEDALUS with Qwen and DeepSeek model pairs, and give every memory to every base agent.

table 2 heuristics transfer across model families
memory generated by
auxiliary / solver
GPT-5.4-miniQwen3.6-35B-A3BDeepSeek-V4-Flash
no memory (MSR)44.3±1.144.8±1.380.6±1.6
GPT-5.4
GPT-5.4-mini
+15.9±1.4+16.1±1.6+8.3±1.8
Qwen3.8-Flash
Qwen3.6-35B-A3B
+16.8±2.4+29.3±1.7+4.9±2.0
DeepSeek-V4-Pro
DeepSeek-V4-Flash
+11.4±3.6+14.5±1.3+3.0±3.0

Hover a cell. ◆ a family using its own memory.

All nine gains are positive. GPT-5.4-mini gains as much from the Qwen memory as from its own (+16.8 vs. +15.9); DeepSeek-V4-Flash gains the least, from an 80.6% baseline that leaves little room. With the GPT-5.4 memory, the seven base agents of fig. 1 all succeed more often, in fewer turns.

06What makes generation work

Which parts of the pipeline make the memory work? A cumulative ablation on AppWorld answers it: we rebuild the pipeline one component at a time, each row adding its component to the previous one.

table 3 cumulative ablation of the generation pipeline

(A) Single Explorer
Heuristics drawn from one 100-turn Explorer trajectory. Exploring alone hurts: the agent does worse than with no memory at all.

(B) Multi-session Explorer
90 sessions of 40 turns, each seeing the heuristics accepted so far. Eighteen times the cost of (A), still below the baseline.

(C) Explorer tasks and Solver traces
The Explorer now proposes tasks, and a heuristic is drawn from every Solver trajectory. +15.8 MSR: what is worth remembering lies in the agent's own attempts.

(D) Solver loop
Heuristics are revised after each failure and kept only after three successes in a row. +4.5 MSR: validation removes the noisy lessons.86.7% of sessions accepted · 183 refinements

(E) Explorer guideline memory
Refinement lessons carried into later sessions: fewer refinements (183 → 161) and 12% cheaper, with MSR within error.91.1% of sessions accepted · 161 refinements

(F) Environment survey (DAEDALUS)
A target distribution of tasks over the environment: 97.8% of sessions end with an accepted heuristic, and generation costs 48% less than (D).97.8% of sessions accepted · 118 refinements

MSR↑pass^5↑gen. $↓
no memory44.3█████████14.9–
(A) Single Explorer35.7███████10.1$5.2
(B) + Multi-session Explorer38.7████████11.3$92.0
(C) + Explorer tasks and Solver traces54.5███████████25.0$46.6
(D) + Solver loop59.0████████████30.4$210.4
(E) + Explorer guideline memory57.9████████████31.5$185.3
(F) + Environment survey (DAEDALUS)60.2████████████32.1$109.7
01/06or ← → keys
AppWorld, 90 sessions for every multi-session row. gen. $: one-time generation cost; the bars show MSR (red: below the no-memory baseline).

Heuristics must come from the Solver's own attempts: exploring alone makes the agent worse (A, B), Solver traces bring the first jump (C) and validating them the second (D). The guidelines and the survey then make generation cheaper rather than better. With fewer refinements, the Solver and the Judge cost 60% less, and the whole run costs half as much as (D).

fig. 5 generation cost by agent role
  • explorer
  • solver
  • judge
  • extractor

07Using the memory

Generating good heuristics is half of the story; the other half is how the agent receives them. Heuristics accepted in different sessions overlap. Consolidated into one memory and injected whole at task start, they are both the best and the cheapest option, 10.7 MSR points above the raw heuristics. Retrieving five heuristics per turn helps less, and BM25, embeddings and a random pick are all within error of each other. Retrieval ranks the heuristics by their resemblance to the agent's current turn, but the heuristic the agent needs often reads nothing like that turn: it warns about a mistake still to come, not about what the agent is doing now.

table 4 consolidating and injecting the memory
mean success ratepass^5 (%)$/run
no memory44.3±1.114.9±2.83.2
top-5 heuristics retrieved per turn, from the consolidated memory
Random51.7±1.519.0±3.04.1
BM2552.0±1.322.0±3.23.8
Qwen3 embeddings54.3±1.624.4±3.33.7
whole memory, re-injected every turn
consolidated memory51.5±1.228.0±3.53.5
whole memory, once at task start
raw heuristics49.5±1.026.2±3.44.9
deduplicated heuristics49.8±2.527.4±3.54.8
consolidated · DAEDALUS60.2±0.932.1±3.63.2
AppWorld, five runs each; the dotted line is the no-memory baseline. Consolidation merges the accepted heuristics in one LLM call, removing redundant or overlapping advice; deduplication only drops the heuristics that restate another. Bold: cheapest.

Retrieving more per turn helps, but never reaches the whole memory at start.

fig. 6 retrieval per turn against the whole memory
  • mean success rate
  • pass^5
  • whole memory at start

08Ranking models with generated tasks

The generated tasks have a use of their own, beyond memory. Environments without training tasks usually lack a test set as well, which makes it hard to compare agents on them. The tasks DAEDALUS generates could fill that role.

To test it, we evaluate nine models, without memory, on two task sets: the 81 tasks accepted during a 90-session generation run on AppWorld, and AppWorld's official test_normal split. If the generated tasks are a good proxy, the models should come out in the same order on both.

fig. 7 generated tasks vs. AppWorld test_normal
● MSR τ = 0.89 (p < .001) · 34/36 pairs agree
○ pass³ τ = 0.89 (p < .001) · 34/36 pairs agree
  • nemotron-3.5-lightning
  • gpt-oss-120b
  • minimax-m2.7
  • gpt-5.4-mini
  • qwen3.6-35b-a3b
  • gemma-4-31b-it
  • mimo-v2.5
  • deepseek-v4-flash
  • gpt-5.6-luna
Each model appears twice: its mean success rate (MSR, filled) and its pass³ (hollow). Hover or tap a model to single it out.
data table
model                         MSR aw     MSR gen    pass³ aw   pass³ gen
────────────────────────────────────────────────────────────────────────
nemotron-3.5-lightning           6.5        13.6         0.2         1.2
gpt-oss-120b                    19.2        39.9         3.2        12.3
minimax-m2.7                    25.1        43.6        10.9        21.0
gpt-5.4-mini                    44.3        66.7        22.2        40.7
qwen3.6-35b-a3b                 44.8        67.9        27.7        46.9
gemma-4-31b-it                  51.0        61.3        32.4        33.3
mimo-v2.5                       52.6        76.0        33.2        51.2
deepseek-v4-flash               60.8        81.5        37.4        61.7
gpt-5.6-luna                    66.2        90.1        49.7        82.7

aw = AppWorld test_normal · gen = DAEDALUS-generated tasks · values in %

The rankings agree. Kendall's τ is 0.89 for mean success rate and 0.89 for pass³ (both p < .001), and 34 of the 36 pairs of models are ordered the same way on both sets.

Absolute scores differ: every model does better on the generated tasks, which were calibrated to GPT-5.4-mini, the solver that generated them. But the order of models is preserved, so tasks generated by DAEDALUS can serve as a proxy for ranking models when no curated test set exists.

09Key takeaways

  • No training tasks, no verifier. DAEDALUS builds an agent's memory from tasks it generates itself.
  • Up to +15.9 points of mean success rate over no memory: the strongest method without training tasks, and competitive with those that learn from curated ones, at little or no extra inference cost.
  • Validation is what works. Heuristics grounded in the agent's own failures and confirmed by its successes help; heuristics from exploration alone hurt.
  • The memory transfers across model families.
  • Give the whole memory at task start. Consolidated and injected once, it beats per-turn retrieval.
  • The generated tasks rank models in the same order as AppWorld's official test set (τ = 0.89).

@Citation

If DAEDALUS, its code or its artefacts are useful to you, please cite:

bibtex
@misc{edy2026daedalus,
  title         = {{DAEDALUS}: Bootstrapping Agent Memory from Self-Generated Tasks},
  author        = {Edy, Antoine and Conti, Max and Xing, Victor and Allard, Marc-Antoine and Benhamdane, Nawfal and Viaud, Gautier},
  year          = {2026},
  eprint        = {2610.08048},
  archivePrefix = {arXiv},
  url           = {https://arxiv.org/abs/2610.08048}
}

AAppendix

three more analyses: the judge, the budget, the auxiliary model

Can the judge be trusted?

The loop rests on an LLM judge. It agrees substantially with each benchmark's own verifier (κ > 0.7), and five calls on the same trajectory agree with each other (κ > 0.8). Its precision beats its recall everywhere: it rejects some successes rather than accepting failures, the safe side for accepting heuristics.

table A1 the LLM judge against the official verifiers
Nprecisionrecallκ benchκ inter
AppWorld168.943.846.807.848
τ²-bench40.889.842.7491.000
AutomationBench70.863.807.729.902
Success counts as positive. κ bench: agreement with the benchmark's verifier; κ inter: agreement between five judge calls on the same trajectory.

How many sessions?

Five sessions already bring more than half of the gain of the 90-session peak, so a useful memory is cheap. Past 90 sessions, success drops while the memory keeps growing.

fig. A1 scaling with the number of sessions
  • mean success rate
  • pass^5
  • heuristics in memory
One generation run, evaluated at checkpoints; error bars are standard errors over five runs.

Does it need a large model?

With GPT-5.4-mini for every auxiliary agent, the solver's own model, DAEDALUS still adds 6.9 MSR points: about half of the gain with GPT-5.4.

table A2 a smaller model for every auxiliary agent
auxiliary modelMSR↑pass^5↑gen. $↓
no memory44.3±1.114.9±2.8–
GPT-5.4-mini51.2±1.620.2±3.1$67.2
GPT-5.460.2±0.932.1±3.6$109.7