Last reviewed: 2026-09-30

Direct answer

Virtual clock testing for coding agent patches should replace any test sleep whose only purpose is to let time pass. Put wall-clock reads and timer scheduling behind one controllable boundary, start each test at a known instant, advance time deliberately, and assert behavior immediately before and at the deadline. A retry scheduled for 1,000 milliseconds should not run after 999 milliseconds and should run after the final millisecond. That pair of assertions catches boundary defects that a long real sleep can miss.

A virtual clock does not remove the need for synchronization. The test must first allow the code under test to schedule its timer, then advance the clock, then let resulting asynchronous work settle. Advancing too early can produce a passing test that never exercised the intended branch. The test should also end with no unexpected timers, tasks, or goroutines left behind.

Different runtimes expose this pattern differently. The Go synctest package documentation describes isolated test bubbles with a fake clock that advances when the contained goroutines are durably blocked. The Node.js test-runner documentation exposes mock timers that can control timer callbacks and the date clock together. Both illustrate the same review principle: elapsed time should be test input, not an uncontrolled environmental dependency.

Keep one narrow integration check for the real scheduler and runtime wiring. Use virtual time for exhaustive unit-level boundaries, retry sequences, expiration rules, and cancellation behavior. Use real time only where the contract being tested is the actual integration with the runtime timer system.

Who this is for

This guide is for maintainers and reviewers evaluating agent-written changes to retry loops, request deadlines, cache expiration, leases, scheduled jobs, debounce logic, or timeout-driven cleanup. It is especially useful when a patch adds sleep, raises a timeout, or repeats a flaky test until it passes.

It is not a replacement for performance measurement or concurrency testing. A virtual clock can prove temporal state transitions quickly, but it cannot establish wall-clock latency, operating-system scheduling behavior, or freedom from data races.

Key takeaways

  • Inventory every time source touched by the changed code: wall-clock reads, elapsed-time measurements, timers, tickers, scheduled callbacks, and retry backoff.
  • Make production code receive a clock or scheduler through a narrow interface rather than reading global time throughout the implementation.
  • Test the instant before a boundary, the boundary itself, and the state after the boundary when that state is meaningful.
  • Let the operation schedule its timer before advancing virtual time, and drain the relevant asynchronous queue afterward.
  • Exercise success, terminal failure, cancellation, and cleanup. A happy-path retry alone is incomplete.
  • Assert that no unexpected timer remains after completion. Leaked timers can affect later tests even when the current assertion passes.
  • Preserve a small real-time integration test, but do not use real sleeps to cover a large state matrix.

Sources checked

  • The Go testing/synctest reference documents isolated bubbles, per-bubble fake clocks, durable blocking, automatic clock advancement, and restrictions around outside I/O.
  • The Node.js test runner reference documents MockTimers, explicit clock ticks, timer cleanup, and coordinated mocking of dates and timers.
  • The Jest timer mocks guide shows timer replacement, advancing by a chosen interval, running pending timers, and the special handling required for recursive timers.
  • The Sinon.JS fake timers guide describes a controllable clock for timer functions, dates, Temporal values, and asynchronous code.

All four pages were refetched successfully for this review. They support the timer-control mechanics described here; project-specific retry policy, expiration semantics, and clock precision must still come from the repository under review.

Contract details to verify

Identify the complete time surface

Start with the changed production path, not the test. Search for direct clock reads, timer construction, retry libraries, expiration comparisons, and background scheduling. A patch is only deterministic if every time source involved in the assertion is controlled. Mocking setTimeout while leaving a separate date read on wall time creates two clocks that can disagree.

Classify each dependency before editing:

  • A wall clock answers what time it is and supports deadlines or expiration timestamps.
  • An elapsed-time clock measures durations and should not depend on calendar adjustments.
  • A scheduler wakes work after a delay or at a deadline.
  • An asynchronous queue runs continuations created by timer callbacks.
  • External I/O can unblock independently of the virtual clock and usually belongs behind a fake at this test layer.

A small injected contract is often enough:

export interface Clock {
  nowMs(): number;
  sleep(ms: number): Promise<void>;
}

The production adapter can delegate to the runtime. The test adapter owns the current instant and the pending sleeper queue. Do not widen this interface with application policy such as retry counts or cache rules; those belong to the component being tested.

Prove the exact boundary

The following Node.js example uses the runtime test context rather than waiting one real second. It proves both sides of the retry boundary and uses the global timer function that the mock controls.

import assert from 'node:assert/strict';
import { test } from 'node:test';

async function retryOnce(run, delayMs) {
  try {
    return await run();
  } catch {
    await new Promise((resolve) => setTimeout(resolve, delayMs));
    return run();
  }
}

test('retries at the exact boundary', async (t) => {
  t.mock.timers.enable({ apis: ['setTimeout'] });
  t.after(() => t.mock.timers.reset());

  let calls = 0;
  const result = retryOnce(async () => {
    calls += 1;
    if (calls === 1) throw new Error('temporary failure');
    return 'ok';
  }, 1000);

  await Promise.resolve();
  assert.equal(calls, 1);

  t.mock.timers.tick(999);
  await Promise.resolve();
  assert.equal(calls, 1);

  t.mock.timers.tick(1);
  assert.equal(await result, 'ok');
  assert.equal(calls, 2);
});

The initial microtask turn matters: it lets the rejected operation reach the timer-scheduling branch. The clock then advances in two steps so the test can distinguish early execution from correct execution. Framework helper names differ, but the sequence remains schedule, assert, advance, settle, and assert again.

Run a concrete happy and error workflow

Use this operator workflow when reviewing a patch:

  1. Record the intended contract in a small matrix. Include the start instant, boundary, expected state immediately before the boundary, expected state at the boundary, maximum attempts, and cancellation outcome.
  2. Run the focused test before changing it. Save the command and result so reviewers can distinguish an existing failure from the patch result.
  3. Replace the real sleep with the repository’s clock seam or the runtime’s supported timer mock. Freeze the wall clock too if the code compares timestamps.
  4. For the happy path, script one temporary failure followed by success. Confirm attempt one runs immediately, no retry occurs before the delay, and attempt two runs exactly at the boundary.
  5. For the error path, script failures through the configured maximum. Advance through each expected delay, assert the final error class, assert the exact attempt count, and verify that no additional retry is scheduled.
  6. Add cancellation if the component accepts a cancellation signal. Cancel while a retry is pending, advance beyond the original deadline, and prove that the operation does not run again.
  7. Restore global timer state and assert that the test adapter has no pending work. Cleanup must run even when an assertion fails.
  8. Run the focused test repeatedly, then run the containing suite. Repetition is evidence of stability, not a substitute for exact assertions.

A minimal command sequence might be:

node --test test/retry.test.js
for run in 1 2 3 4 5; do
  node --test test/retry.test.js || exit 1
done

The error-path result should be deterministic: the same attempt count, final error class, virtual elapsed time, and pending-timer count on every run. If those values vary, the patch still has an uncontrolled dependency.

Keep logs useful and sanitized

Test evidence should describe the temporal transition without copying request bodies, response bodies, environment values, headers, or arbitrary exception text. Emit an allowlisted record such as:

event: virtual_clock_check
test_case: retry_after_temporary_failure
clock_mode: virtual
clock_start: 2026-01-01T00:00:00Z
advanced_ms: 1000
boundary: at_deadline
attempt_count: 2
pending_timer_count: 0
result: passed
error_class: none

For an error path, set result to failed_as_expected and record a stable error class rather than the full message. These fields let a reviewer confirm which boundary ran and whether cleanup completed without exposing application data.

Preserve a narrow real-time check

The virtual suite should own the state matrix. A separate integration test can confirm that the production clock adapter delegates to the runtime and that one scheduled callback is eventually observed. Give that check a generous upper bound and avoid asserting an exact wall-clock duration. If it fails, the failure should identify runtime integration rather than obscure the retry policy already proved under virtual time.

Failure modes

  • Only the timer is mocked. Code reads the real date while callbacks follow virtual time, so expiration and scheduling disagree. Control every clock used by the decision.
  • Time advances before the timer exists. The test calls tick immediately after starting asynchronous work, but the retry branch has not scheduled anything. Let the scheduling path settle first and assert the pending state.
  • The test skips the pre-boundary assertion. Advancing directly to the deadline proves that work eventually runs but does not prove it avoided running early.
  • Recursive timers are exhausted indiscriminately. The Jest guide notes that running all recursive timers can form an endless loop. Run only the pending generation or advance a bounded interval, then assert what was newly scheduled.
  • External I/O sits inside an isolated virtual-time test. The Go reference distinguishes durable in-bubble blocking from operations such as network I/O that can be released externally. Replace that I/O with an in-process fake or test it in a separate integration layer.
  • Timer callbacks run but continuations do not settle. Advancing the timer queue may not finish promise or task continuations. Use the framework’s asynchronous timer helper or explicitly await the operation before reading final state.
  • Cleanup is omitted. A pending interval or un-restored global clock contaminates later tests and makes results order-dependent.
  • Units are converted implicitly. A value interpreted as seconds in one layer and milliseconds in another can make a virtual test pass at the wrong boundary. Name units in interfaces, variables, assertions, and log fields.
  • A higher timeout is presented as a fix. Increasing a real delay can reduce observed failures while making the suite slower and leaving the underlying race intact.
  • Virtual tests replace every integration check. The policy logic may be correct while the production adapter is wired to the wrong runtime function. Keep one focused real-scheduler check.

FAQ

Does virtual time prove that code is free of races?

No. It makes temporal state transitions repeatable, but it does not explore every thread or goroutine interleaving. Treat clock control and race detection as separate evidence. The race-condition testing guide covers the concurrency side.

Should every sleep disappear from the repository?

No. Production code may legitimately sleep, delay, or schedule work. The review target is uncontrolled waiting in tests and direct time access that prevents deterministic substitution. Even in tests, a small real-time integration check can be appropriate when it verifies the runtime adapter itself.

What boundaries should a reviewer require?

At minimum, require an assertion immediately before the deadline and another at the deadline. Add after-deadline behavior when the contract permits repeated scheduling, expiration cleanup, or grace periods. Also test zero, negative, or maximum durations if the repository contract accepts them.

What if the framework cannot report pending timers?

Wrap scheduling in a project-owned adapter that tracks outstanding handles in test mode. The assertion should be part of the adapter’s test contract, not an inspection of unrelated runtime internals.

How should an agent be instructed to change a flaky timing test?

Specify the temporal contract and evidence, not merely the desired tool. Ask for a controllable clock, before-and-at-boundary assertions, terminal failure and cancellation cases, cleanup verification, and the exact focused test command. Do not ask the agent simply to make the test pass faster.

Reader next step

Choose one changed test that currently waits for wall time. Write a four-row matrix for initial state, immediately before the deadline, at the deadline, and terminal error. Then inject or enable a controllable clock, implement those assertions, verify zero pending work, and run the focused test repeatedly before the full suite.

Use the flaky-test detection workflow to decide whether remaining intermittent failures belong in quarantine or indicate a regression. If concurrent state is still involved after time is controlled, continue with the race-condition testing guide rather than adding another sleep.