2026-10-11 17:14 UTC

Trigora's Omar Abdelrahman claims its demonstrated Transparent Continuation Checkpointing prototype restores durable executions from live continuations rather than replaying history, potentially decoupling long-running agent recovery costs from accumulated execution history.

state: watchingheat: lowuncertainty: mediumconvergesscott: mediumagent-harnesses durable-execution long-running-agentsOmar AbdelrahmanTrigora

What is this?

Trigora’s research page describes Transparent Continuation Checkpointing (TCC) as a durable-execution architecture that captures resumable program state at durable boundaries. The case attributes to Omar Abdelrahman a prototype that restores executions from live continuations rather than replaying history, potentially reducing history-dependent recovery costs for long-running agents. The supplied web snippet establishes the checkpointing architecture, but does not verify Abdelrahman’s role, the demonstration, the reported 0.6–0.9 ms recovery time, or independence from accumulated history.

Why it matters to Scott

TCC converges with Scott’s checkpoint-and-resume direction and offers a runtime-level mechanism worth evaluating for Proposal Compiler’s persisted job recovery, rather than merely repeating the case for durable agents. Continuation restoration is distinct from his compressed agent handovers, and the claimed history-independent recovery and 0.6–0.9 ms timing remain unverified; the supplied radar hits track related runtimes, not this development.
ip:framework.long-running-agentsip:concept.checkpoint-disciplinedev:concept.resumable-agent-job-control-planedev:project.proposalradar:concept.durable-agentsradar:pi-agentharness-durable-runtimeradar:ava-durable-replayable-sessions
queries asked of Scott's wikis
  • agent harness crash recovery and durable execution
  • checkpoint resume versus event history replay
  • long-running agent recovery cost scaling
  • agent execution state persistence and serialization
  • durable boundaries side effects and retry correctness

Measured heat

now 0 pts/hpeak 3 pts/hcomments 0/hpeers p14momentum: steady2 platformsage 794h
points/hour across evidence · reading as of 2026-10-12 02:59:37.977291+11:00 · deterministic, not a model opinion

How the heat travelled

09-08 14:00⭐ origin echo-reconstructedIntroduces Transparent Continuation Checkpointing and links a technical paper and controlled demonstration; reports 0.6–0.9 ms recovery with
Omar Abdelrahman on blog (echo) · attributed from hn.story.49635621
—
09-09 22:41first on hacker news · published · +32.7hDurable execution without history replay
hypervs
—
09-09 22:41amplified on hacker news 👑hn.story.49635621
hypervs
peak 41 · 32 comments · 86% of case engagement
10-07 15:44amplified on hacker newshn.story.49994446
hypervs
peak 10 · 2 comments · 14% of case engagement
09-13 05:21our radar first saw it · +111.3hdiscovery anchor: hn.story.49635621—
pace: p65 vs 519 stories at the 720h mark (now 794h old) — ahead of runway-gwm-worlds-2 (1.0x), behind geiger-local-agent-access-inventory (1.0x)

Evidence (3) — ⭐ canonical anchor

sourceobjectauthorscorecomments
🟧 hnDurable execution without history replay
Retrieved article excerpt

Open article · Retrieved 2026-09-13T05:21:53.367997+00:00

[Blog](https://trigora.dev/blog)

# Durable execution without history replay

[Omar Abdelrahman](https://trigora.dev/cdn-cgi/l/email-protection#1d72707c6f5d696f747a726f7c3379786b) · 9 September 2026

Most durable execution systems recover by replaying retained execution history. After a worker fails, a fresh worker loads the history and re-executes the program until it reconstructs the current position.

This is a useful model. It provides durable progress while allowing workers to remain ephemeral. But it also makes accumulated history part of the recovery path.

That tradeoff becomes more noticeable for programs that operate for hours or days, call many tools, wait for external events, create child executions, and change direction dynamically. Long-running agents increasingly have this shape.

I’ve built and evaluated a different recovery primitive: checkpointing the program continuation instead of reconstructing it from history.

## Transparent Continuation Checkpointing

I call the approach Transparent Continuation Checkpointing, or TCC.

At durable boundaries, the compiler and runtime capture the live continuation: the control state required for the program to continue from its current position. When execution resumes after a failure, the runtime loads the committed continuation and restores the program directly.

  History replay     reconstruct TCC      resume  

History replay reconstructs the current position. TCC restores the committed continuation
and resumes.

 

The distinction is:

**History replay**

Load retained history → re-execute the prefix → reconstruct the current position

**TCC**

Load committed continuation → restore live execution state → resume

External effects remain explicit durable operations. Completed durable work is not repeated after recovery, and unsupported language constructs fail during compilation rather than producing ambiguous runtime behaviour.

The current prototype supports durable effects, external waits and events, child executions, cancellation, structured concurrency, and crash recovery.

## What changes

TCC does not make recovery constant-time. Recovery remains sensitive to the size and structure of the live continuation.

The intended change is in what recovery depends on.

With replay, recovery is influenced by the execution history retained to reconstruct the current position. With TCC, recovery is influenced primarily by the state the program still needs.

A program that has performed ten thousand operations but retains a small live continuation should not necessarily become harder to recover simply because its past is long.

## Preliminary evaluation

I ran a controlled comparison in which live continuation state remained approximately fixed while durable-boundary depth increased from 10 to 1,000.

  Recovery latency vs prior durable-boundary depth    1ms     10ms     100ms     1s    10  100  1000              
durable-boundary depth
  

TCC recovery
  
Temporal reconstruction
 Controlled evaluation · ~4 KB live state

 

In that evaluation, TCC recovery remained between approximately 0.6 and 0.9 milliseconds. Fresh-worker replay reconstruction in the evaluated Temporal baseline increased from approximately 61 milliseconds to 1.7 seconds.

Worker creation was excluded, the live state was approximately 4 KB, and these results should not be interpreted as a general production-speedup claim. They demonstrate a difference in recovery scaling under the tested conditions, not that every TCC workload will outperform every replay-based system.

[Methodology and limitations](https://trigora.dev/research/recovery-vs-history)

I have also exercised the execution semantics across 50,000 generated cases, with no observed semantic failures in the evaluated subset.

## What remains difficult

Turning the prototype into production infrastructure still involves substantial work:

- Portable continuation representation
- Program and checkpoint versioning
- Efficient handling of larger live states
- Durable storage and commit protocols
- Operational observability
- Compatibility across language frontends
- Framework integrations
- Long-running correctness and failure testing

There are also design questions around checkpoint retention, branching from previous continuations, migration between runtime versions, and how much of the execution representation should remain stable across languages.

I’m building Trigora around this model, initially for long-running AI agents. The broader question is whether continuation-based recovery can provide a better execution substrate for dynamic, long-lived software.

The architecture, semantics, benchmark setup, and current limitations are described in more detail in the [technical paper](https://trigora.dev/research/whitepaper). You can also [see continuation recovery in action](https://demo.trigora.dev) in a controlled demonstration of the current TCC compiler/runtime.

I’d be particularly interested in criticism from people who have worked on workflow engines, compilers, checkpointing systems, or distributed runtimes.
hypervs4132
🟧 echo.blog ⭐Introduces Transparent Continuation Checkpointing and links a technical paper and controlled demonstration; reports 0.6–0.9 ms recovery withOmar Abdelrahman——
🟧 hnShow HN: Trigora – durable execution without history replayhypervs102

Interpretation history

Decision trace