Skip to content

[Showcase] Emergent 3-phase curriculum from an intrinsically-motivated agent #33

Description

@augo-augo

Hi Crafter team,

Thank you for building this benchmark, it’s been incredible to work with.

I wanted to share something unexpected that emerged while training an intrinsically-motivated agent in Crafter. The agent (DTC - Dual-Timescale Competence) learns without external rewards, using only a competence-based intrinsic signal.

What emerged was a clear three-phase developmental pattern:

Phase 1 (0–80k steps): Skill Mastery
The agent masters basics like wake_up and collect_wood, practicing them heavily.

Phase 2 (80k–140k steps): Boredom Trough
My competence reward (difference between slow/fast novelty EMAs) drops to zero as skills are mastered. The agent gracefully deprecates these skills and stops practicing them—you can see wake_up activity drop off completely.

Phase 3 (140k+ steps): Exploration Cascade
The “boredom” state triggers adaptive exploration (boosted entropy via CognitiveWaveController), leading to a cascade of compositional discoveries:

  • collect_sapling
  • place_plant
  • make_wood_sword
  • defeat_zombie

I haven’t seen this pattern documented before, but it might speak to the richness of skill composition in Crafter’s design.

Links:

Thanks for your work on this!

Augo

Activity

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    No labels
    No labels

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions