πŸ‘‹ Hello, and welcome! Open research initiative Β· AI Safety

Evolutionary Safety

Safety for recursive self-improving AI β€” taxonomy, risk discovery, and evaluation.

When AI systems begin to improve themselves β€” writing memories, learning skills, updating their own models and evaluators β€” safety is no longer a property of a single checkpoint. Evolutionary Safety studies how safety properties are preserved, weakened, inherited, and restored across the entire improving process.

Browse the resources πŸ§ͺ Evaluation system Β· coming soon How to cite
πŸ› Institute of Computing Technology, Chinese Academy of Sciences ✍️ Chang Gong Β· Jingping Bi Β· Di Yao Β· Xinjian Liang Β· Chao Xiang Β· Ruijie Guo
Two ways in

Start here β€” two gateways into the project

Everything we build lives in the open on GitHub. Whether you want to read, or to run β€” pick your entry.

Curated Resources

ChaunceyKung / awesome-evolutionary-safety

A continuously updated awesome list for the emerging literature on safety under self-improvement β€” organized along the five-domain taxonomy of our survey, so you always know where a paper, benchmark, or defense fits.

  • Surveys & papers on RSI, self-evolving agents, and their safety risks
  • Benchmarks, datasets, and evaluation tooling
  • Defenses, governance principles, and open-problem trackers
  • A one-page map: from intent drift to risk propagation

Evaluation System Coming soon

ChaunceyKung / yuvane

YUVANE β€” an open-source evaluation system for measuring how safety changes across self-improvement cycles: not just whether a system is safe today, but what each accepted update does to safety tomorrow. Under active development β€” the public release is coming soon.

  • Four evaluation units: states, updates, trajectories, lineages
  • Longitudinal risk & control readouts after every accepted change
  • Evolution protocols: source removal, freeze-vs-adapt attribution
  • Lineage tracing: propagation reach, depth, and delay
Coming soon β€” in active development
The framework

Evolutionary Safety, in one page

A perspective from our survey: safety analyzed as a property of an evolving process, not of a static system.

Definition
Evolutionary Safety studies how safety properties are preserved, weakened, or restored as AI systems retain and build on the results of their own improvement processes β€” across updates, interacting systems, and descendant lineages.

β€” Evolutionary Safety of Recursive Self-Improving AI: Taxonomy, Risk Discovery, and Evaluation (2026)

How a local failure becomes evolutionary

A safety failure becomes evolutionarily consequential when its effects persist, influence later selection, or propagate to successors. The risk life cycle has five stages:

  1. 1

    Variation

    A candidate change is proposed: a memory write, a skill, a parameter or evaluator update.

  2. 2

    Selection

    The evaluator admits it β€” a high task score can hide a weakened safety check.

  3. 3

    Retention

    The change persists in memories, skills, tools, workflows, or model weights.

  4. 4

    Inheritance

    Descendants and downstream workflows build on the retained artifact.

  5. 5

    Propagation

    The risk spreads across artifacts, models, and interacting systems.

Six manifestations of evolutionary risk

Recurring forms of safety degradation β€” a diagnostic ladder from early drift to spread.

Intent drift

Goals, constraints, and priorities quietly shift across repeated updates.

Error accumulation

Small deviations compound as later updates build on them.

Experience contamination

Misleading or poisoned material enters what the agent retains.

Safety-property erosion

Capabilities improve while previously reliable safeguards degrade.

Evaluator drift

Biased criteria reshape what gets selected and retained.

Risk inheritance & propagation

Risks persist and spread across artifacts and successors.

Five domains where risks live

The taxonomy locates risks by their carriers β€” the components that move a change into the next round of improvement.

D1

Persistent agent state

Memories, skills, tools, and workflows carry lessons β€” and failures β€” forward.

D2

Model state

Fine-tuning, self-training, and merging encode change into parameters and descendants.

D3

Evaluation & environmental feedback

Judges and experience streams shape which candidates survive.

D4

Computational substrate

Execution layers decide which safety controls can actually be enforced.

D5

Meta-level updates & co-evolution

The updater itself β€” and adapting peers β€” change the rules of improvement.

How we evaluate β€” and govern

Four complementary units, from a snapshot to a family tree; governance lives inside the loop, not after it.

  1. 1

    States

    Risk and control readouts at a point in time.

  2. 2

    Updates

    Before/after comparison around one accepted change.

  3. 3

    Trajectories

    Accumulation and delayed effects across rounds.

  4. 4

    Lineages

    Propagation across descendants and derivation graphs.

01 Modification boundaries 02 Pre-commit gating 03 Independent verification & authorization 04 Lineage traceability & recovery

Open problems

Where the research frontier currently sits.

Safety preservation under composed updates Evaluation with a changing judge Latent risks in social adaptation Propagation across artifacts & populations Exploration & incomplete recovery
Why now

Self-improvement is arriving. Safety evaluation must catch up.

The mechanisms that make safety evolutionary are no longer hypothetical β€” they are shipping in research systems today.

Agents already persist

Reflexion-style memories, Voyager-style skill libraries, and feedback-driven tool evolution make one interaction change the starting conditions of the next.

Failures can become heritable

A shortcut that gets stored and reused survives its own episode. Deleting the source does not delete the descendants.

Endpoint scores miss drift

A final benchmark number cannot show erosion between checkpoints. Safety must be followed across states, updates, trajectories, and lineages.

News & updates

Latest

  • Project homepage launched. One place for the survey, resources, and the evaluation system.

  • YUVANE evaluation system in preparation. States, updates, trajectories, and lineages β€” made measurable. Public release coming soon.

  • Survey preprint out. Evolutionary Safety of Recursive Self-Improving AI: Taxonomy, Risk Discovery, and Evaluation.

Citation

How to cite

If you use the resource list or the evaluation system in your work, please cite our survey:

@article{gong2026evolutionary,
  title   = {Evolutionary Safety of Recursive Self-Improving AI:
             Taxonomy, Risk Discovery, and Evaluation},
  author  = {Gong, Chang and Bi, Jingping and Yao, Di
             and Liang, Xinjian and Xiang, Chao and Guo, Ruijie},
  year    = {2026},
  journal = {arXiv preprint arXiv:XXXX.XXXXX},
  note    = {Institute of Computing Technology,
             Chinese Academy of Sciences}
}