Clear, practical technology insights BSOD Code Lookup · Windows Error Code Lookup · Wi-Fi Troubleshooting · PC Troubleshooting Checklist

Memento-Skills: How AI Agents Improve Skills Without Retraining the LLM

How Memento-Skills updates an external skill library, what its GAIA and HLE results measure, and the safeguards needed before using self-modifying agents.

Table of Contents

Memento-Skills is an AI agent system that improves reusable instructions and supporting code while keeping the underlying language model's parameters unchanged. Its learning happens in an external skill library: the agent retrieves a skill, tries it, examines the result and proposes an update or a new skill.

That distinction matters. The system can adapt its workflow without fine-tuning the base model, but it still needs execution, feedback, evaluation and maintenance. “No LLM retraining” does not mean improvement is free, automatic or guaranteed.

What changes when the model stays frozen?

A conventional model does not permanently learn from every task simply because the conversation contains new information. An agent can nevertheless change its behavior by loading updated instructions, retrieving relevant knowledge or calling different tools.

Memento-Skills makes reusable skills the persistent part of that process. A skill can contain a description of when it applies, instructions for carrying out a task, and code or supporting resources. Instead of keeping only a transcript, the system tries to retain a procedure it can use again.

This is one approach within the wider range of AI agent frameworks. It should not be described as an automatic upgrade built into every existing assistant or as a replacement for all model training.

How the skill learning loop works

  1. Read: select relevant skills for the current task and context.
  2. Execute: use the selected instructions and tools to attempt the task.
  3. Reflect: inspect the outcome and identify whether a skill, an assumption or another part of the workflow caused a failure.
  4. Write: revise an existing skill or author a missing capability, then validate the candidate before retaining it.
Memento-Skills: A new framework that helps AI write skills automatically without retraining the model. Picture 1
The framework organizes adaptation around reusable skills and feedback.

The research uses a behavior-trainable router rather than treating similar wording as sufficient evidence that a skill is useful. For example, a billing question and a refund operation may share vocabulary but need different procedures. The router's job is to select useful behavior, not merely a document containing matching words.

“Without retraining” refers to the base LLM. It does not mean that the routing and memory system have no learning machinery. The paper discusses memory-based reinforcement learning and evolving prompts and skills.

What the reported benchmark results show

The authors' Memento-Skills paper reports the following test results for its research configuration. Its comparison disables skill optimization while retaining retrieval, execution and feedback collection.

BenchmarkComparison without skill optimizationFull systemDifference
GAIA52.3%66.0%13.7 percentage points
Humanity's Last Exam17.9%38.7%20.8 percentage points
Memento-Skills: A new framework that helps AI write skills automatically without retraining the model. Picture 3
Reported benchmark gains apply to the study's evaluation setup, not every deployment.

The paper identifies Gemini-3.1-Flash as the underlying model. Its HLE experiment uses a selected training and test split, so these figures should not be presented as a directly interchangeable score for every public HLE evaluation.

Starting from five basic skills, the researchers report libraries of 41 skills for GAIA and 235 for HLE. More skills do not necessarily mean better performance: the paper also notes limited reuse across dissimilar GAIA tasks. A learned procedure helps most when later tasks require related behavior.

These results measure task accuracy. They do not establish that an enterprise rollout will be faster, cheaper or safer, or that the agent can continually improve without oversight.

Testing and rollback remain essential

The project repository distinguishes its public runtime from the research configuration. It describes guarded skill repair with isolated changes, validation gates, audit history and rollback. Teams should check the specific version they evaluate rather than assuming all paper components are reproduced identically.

Memento-Skills: A new framework that helps AI write skills automatically without retraining the model. Picture 2
Testing a proposed skill change is a control step, not proof that every future use is safe.

An agent that writes both a fix and its test may miss the same edge case twice. Keep independent acceptance cases, review changes to executable code and preserve a known working version. Instructions stored in a skill must not grant new permissions by themselves.

How to evaluate it for a business workflow

A sensible pilot uses repeated, verifiable tasks, such as extracting fields from a consistent document type or preparing a report from approved inputs. Compare a fixed skill library with an evolving one on the same held-out tasks.

  • Measure correct results, total runtime, tool costs and human correction effort.
  • Keep learning examples separate from evaluation examples.
  • Limit file, network and account access to what the pilot requires.
  • Review proposed skill changes before they affect consequential operations.
  • Track regressions on older tasks as well as improvements on new ones.

Developers implementing their own system can also review how to add Agent Skills support. The practical value comes from reliably reusing a better procedure, not from allowing unrestricted self-modification.

Discussion

Reader Comments 0

Sign in with email or Google to join the discussion.