Skip to content

extropy estimate assumes early convergence, ignores timeline events #112

Description

@DeveshParagiri

Bug

extropy estimate predicts ~3 effective timesteps for a 6-timestep evolving scenario, but the simulation will actually run all 6 because early convergence is auto-disabled when future timeline events exist.

Evidence

ASI scenario: 5,000 agents, 6 monthly timesteps, each with a timeline event.

Estimated LLM Calls
                      Calls      Input Tok     Output Tok
  Pass 1              5,122    ~11,104,496     ~1,024,400
  ...
  Total:   $5.93

Estimate says "Effective timesteps: ~3 (early stop at ~100% exposure)" — but the simulation engine disables early convergence when allow_early_convergence is None (auto) and future timeline events exist (stopping.py:366). All 6 timesteps will run, so actual cost is ~2x the estimate.

Expected Behavior

Estimate should apply the same auto-convergence logic as the simulation engine: if the scenario has timeline events at future timesteps, assume all timesteps will run.

Activity

  1. DeveshParagiri commented on Feb 18, 2026

    @DeveshParagiri
    CollaboratorAuthor

    Deeper analysis: estimator is 4-5x off

    The estimator undercounts for three compounding reasons:

    1. Ignores re_reasoning_intensity

    The ASI scenario has re_reasoning_intensity: extreme on timesteps 1, 2, 3, and 6. When extreme, the engine forces ALL aware agents to re-reason (engine.py:656-660). The estimator doesn't model this at all — it uses a flat 2% re-reasoning rate (estimator.py:214).

    2. Ignores early convergence auto-disable

    As noted in the original issue, the estimator assumes early stop at ~3 timesteps. The engine runs all 6 because allow_early_convergence: null + future timeline events = disabled.

    3. Doesn't model conversations

    Medium/high fidelity adds 4-8 LLM calls per conversation (all fast model). The estimator counts zero conversation calls.

    Actual vs estimated for ASI (5,000 agents × 6 timesteps)

    Estimator Actual (projected)
    Effective timesteps 3 6
    Reasoning calls ~5,100 ~21,000-22,000
    Total LLM calls ~10,200 ~42,000+
    Cost $5.93 $25-35

    Per-timestep breakdown (projected)

    Timestep Intensity Reasoning calls Why
    1 extreme ~4,000 Seed exposure (broadcast)
    2 extreme ~5,500 Remaining exposed + ALL re-reason
    3 extreme ~5,000 ALL aware agents forced
    4 high ~300 Multi-touch only
    5 high ~1,500 Some re-reasoning
    6 extreme ~5,000 ALL aware agents forced

    What needs fixing

    The estimator needs to:

    1. Read re_reasoning_intensity from each timeline event and model forced re-reasoning (extreme = all aware, high = fraction)
    2. Apply the same early convergence auto-disable logic as the engine
    3. Account for conversation calls based on fidelity setting
  2. DeveshParagiri commented on Feb 18, 2026

    @DeveshParagiri
    CollaboratorAuthor

    Example: estimator vs reality for ASI scenario (5,000 agents × 6 timesteps)

    What the estimator says

    Effective timesteps: ~3 (early stop at ~100% exposure)
    Total calls: ~10,200
    Cost: $5.93
    

    What actually runs (projected)

    The estimator misses three things: (1) re_reasoning_intensity: extreme forces ALL aware agents to re-reason on timesteps 1/2/3/6, (2) early convergence is auto-disabled because every timestep has a timeline event, (3) conversations are not counted.

    Reasoning calls per timestep:

    Timestep Intensity Reasoning calls Why
    1 extreme ~4,000 Seed exposure (broadcast)
    2 extreme ~5,500 Remaining exposed + ALL re-reason
    3 extreme ~5,000 ALL aware forced
    4 high ~300 Multi-touch only
    5 high ~1,500 Some re-reasoning
    6 extreme ~5,000 ALL aware forced
    Total ~21,000

    LLM calls breakdown by fidelity (gpt-5-mini @ $0.25/$2.00 per MTok):

                            MEDIUM FIDELITY          HIGH FIDELITY
    ─────────────────────────────────────────────────────────────
    Reasoning events:       ~21,000                  ~21,000
    Pass 1 calls:           21,000                   21,000
    Pass 2 calls:           21,000                   42,000 (2× for public stmt)
    Conversation calls:     ~21,000 × 15% × 1 × 4   ~21,000 × 15% × 2 × 6
                            = ~12,600                = ~37,800
    
    Total LLM calls:        ~54,600                  ~100,800
    

    Cost:

    Medium High
    Pass 1 (21K × ~2.2K in / 200 out) ~$20 ~$20
    Pass 2 (21-42K × ~300 in / 70 out) ~$4.5 ~$9
    Conversations (12-38K × ~800 in / 150 out) ~$6 ~$18
    Total ~$30 ~$47

    Time at 1000 RPM (Azure gpt-5-mini):

    Medium High
    Total calls ~55K ~101K
    At 1000 RPM ~55 min ~101 min
    With burst headroom ~45-55 min ~80-100 min

    Estimator vs actual:

    Estimator Actual (medium) Actual (high)
    Calls ~10K ~55K ~101K
    Cost $5.93 ~$30 ~$47
    Time N/A ~50 min ~90 min

    Conversation % is a guess (15% of agents request talk_to). For a scenario like ASI where everyone has strong opinions, could be 30%+ — which would double the conversation line items.

  3. RandomOscillations commented on Feb 24, 2026

    @RandomOscillations
    Collaborator

    Verification update (2026-02-24): issue still reproduces in current code.

    Current estimator logic remains timeline-unaware:

    • extropy/simulation/estimator.py:235-238 applies a generic exposure early-stop.
    • It does not mirror simulation runtime gating that uses future timeline events / allow_early_convergence policy.

    Given this mismatch, the CLI estimate surface has been temporarily stubbed to prevent misleading pre-run numbers while preserving estimator internals for later parity work.

    Temporary CLI behavior:

    • estimate is hidden from top-level help.
    • Direct invocation returns a temporary-disabled message referencing this issue.

    Keeping this issue open until estimator parity is implemented and tested against runtime stop behavior.

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    No labels
    No labels

    Type

    No type

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions