Skip to content

About

Explainer for the Task Interrupt Timing API: PerformanceObserver entry to detect main-thread jank with configurable thresholds and point-in-time stack trace.

Resources

Code of conduct

Contributing

Stars

1 star

Watchers

2 watching

Forks

Latest commit

 

History

9 Commits

Folders and files

NameName
Last commit message
Last commit date
 
 
 
 
 
 
 
 
 
 
 
 

Repository files navigation

Explainer for the Task Interrupt Timing API

This proposal is an early design sketch by [Chrome Webium] to describe the problem below and solicit feedback on the proposed solution. It has not been approved to ship in Chrome.

Proponents

  • Eriko Kurimoto (@elkurin)

Participate

Table of Contents [if the explainer is longer than one printed page]

Introduction

Modern web applications strive for smooth, responsive user interfaces. However, developers often struggle to identify and eliminate "jank"—micro-stutters caused by JavaScript tasks hogging the main thread.

While tools exist to measure long frames or very long tasks (>50ms), developers lack a programmatic way to detect and diagnose shorter, highly-contended tasks (e.g., 5ms - 15ms) that consistently burn CPU cycles and delay responsiveness. Furthermore, existing APIs provide post-mortem, aggregated script attribution, which often fails to point developers to the exact line of code that was blocking the thread at the critical moment.

The Task Interrupt Timing API proposes a new PerformanceObserver entry type that allows developers to set a custom, low-duration threshold. When a task exceeds this threshold, the browser triggers a mid-execution interrupt to capture a precise, point-in-time stack trace of the offending script.

Goals

  • Provide a programmatic way to monitor individual JavaScript task execution times against custom, low-duration thresholds (e.g., 5ms or 10ms).
  • Capture and expose precise, point-in-time JavaScript stack traces when a task exceeds the configured threshold.
  • Ensure the detection and observation mechanisms have near-zero performance overhead on the main thread's critical path.

Non-goals

  • Measuring total frame duration: This API is explicitly designed to measure isolated JavaScript tasks, not the entire rendering pipeline or layout/paint times. (The Long Animation Frames API already solves this).
  • Replacing DevTools Profilers: This API is for programmatic, lightweight monitoring in the wild, not for capturing full continuous CPU profiles.

Use cases

This proposal is primarily driven by the performance requirements of the Webium Product project, which needs strict, programmatic control over main-thread responsiveness to deliver a stutter-free user experience. Webium's requirements highlight a broader challenge faced by complex, highly interactive web applications.

Use case 1: Diagnosing micro-stutters in the Webium Product

In the Webium Product, even small tasks (e.g., 5-15ms) can cause noticeable micro-stutters if they occur during critical rendering or interaction phases. The project needs to enforce strict sub-50ms performance budgets and understand exactly which JavaScript function is blocking the thread at the moment the budget is breached.

Current post-mortem tools only provide aggregated script attribution after a long frame finishes, which is insufficient for Webium's diagnostic needs. By capturing a precise, point-in-time stack trace exactly when a short threshold is exceeded, the Webium team (and developers of similar complex apps) can pinpoint and eliminate the exact source of main-thread contention in the wild.

Use case 2: Real User Monitoring (RUM) and Telemetry

Performance analytics providers (e.g., third-party RUM scripts) currently struggle to report actionable main-thread contention data from actual users. While they can report that a user experienced a long frame, they cannot programmatically capture the offending JavaScript stack trace to send back to the server. This API would allow telemetry scripts to automatically capture and aggregate point-in-time stack traces for tasks that violate performance SLAs in production, giving site owners actionable data to fix jank.

Potential Solution

We propose introducing a new PerformanceObserver entry type: task-interrupt.

Developers can configure the observer with a custom durationThreshold. When a JavaScript task begins on the main thread, a background monitor starts a timer. If the task continues executing past the durationThreshold, the browser requests an immediate interrupt from the JavaScript engine.

At the next safe execution point, the engine captures the current execution stack trace. Once the task finally completes, a PerformanceTaskInterruptTiming entry is dispatched to the observer.

[Exposed=Window]
interface PerformanceTaskInterruptTiming : PerformanceEntry {
    // entryType will be "task-interrupt"
    // startTime and duration are inherited from PerformanceEntry
    
    // The captured stack trace at the exact moment the threshold was crossed.
    readonly attribute FrozenArray<DOMString> stackTrace;
};

How this solution would solve the use cases

Developers can define a strict budget (e.g., 10ms) and automatically log exact stack traces to their telemetry systems when the budget is violated.

// 1. Create a PerformanceObserver to listen for task interrupts
const observer = new PerformanceObserver((list) => {
  for (const entry of list.getEntries()) {
    console.warn(`Janky task detected. Total duration: ${entry.duration}ms`);
    
    // The API exposes the exact point-in-time stack trace captured when 
    // the threshold was crossed.
    if (entry.stackTrace) {
      console.warn(`Stack trace captured at threshold:\n${entry.stackTrace.join('\n')}`);
    }
  }
});

// 2. Start observing tasks with a custom threshold
observer.observe({ 
  type: 'task-interrupt', 
  durationThreshold: 10 
});

For both strict performance budgeting (like the Webium Product) and Real User Monitoring (RUM), developers can deploy the PerformanceObserver shown above.

When a custom budget (e.g., 10ms) is violated, the stackTrace is synchronously captured by the browser engine and surfaced in the PerformanceTaskInterruptTiming entry. The developer's script can then read this entry and beacon the exact function name and line number back to their telemetry backend, completely eliminating the guesswork in diagnosing production jank.

Detailed design discussion

Mid-execution interrupts vs. Post-mortem aggregation

A core design choice was whether to collect script attribution data continuously during the task (aggregation) or to trigger a single interrupt. We chose the interrupt-driven model because it provides a precise "point-in-time" snapshot of what was blocking the thread exactly when the threshold was breached, resulting in a clearer signal for developers compared to aggregated execution times.

Lock-free task observation

Because the browser must observe every single task on the main thread to measure its duration, introducing locks or thread-synchronization to communicate with a background watchdog thread would severely impact overall browser performance. This design relies on a lock-free, relaxed atomic write on the main thread that the background thread polls, ensuring zero overhead on the critical path.

Considered alternatives

Long Animation Frames (LoAF) API

The Long Animation Frames API is a fantastic tool for measuring responsiveness, but it is fundamentally unsuited for this specific use case for two reasons:

  1. Frames vs. Tasks: LoAF is tied to a 50ms frame rendering threshold. We need to measure individual tasks with a configurable, much lower threshold to catch thread-hogging before a frame is even dropped.
  2. Aggregated Attribution vs. Point-in-Time: LoAF uses a post-mortem, aggregated script attribution model (e.g., "Function X took 40ms total over the course of the frame"). Our design relies on triggering a mid-execution interrupt to capture a precise stack trace.

Long Tasks API

The Long Tasks API measures tasks rather than frames, which aligns closer to our goal. However, its 50ms threshold is hardcoded into the spec. Extending it to accept a custom 5ms threshold and fundamentally changing its attribution model to use mid-execution stack traces would break the existing semantics of the API. Introducing a distinct task-interrupt entry type provides a cleaner separation of concerns.

Security and Privacy Considerations

Exposing raw JavaScript stack traces to the web platform introduces significant privacy and security risks. Specifically, a malicious site could embed a cross-origin resource (such as a third-party script) and use this API to observe its internal execution flow. By reading the function names in the stack trace, the attacker could infer sensitive user state (e.g., inferring authentication status based on which code paths are executing).

To mitigate this cross-origin data leakage while keeping the API usable on general websites (which often rely on ads and analytics), we propose stack truncation at cross-origin boundaries.

Instead of strictly requiring Cross-Origin Isolation (COOP/COEP) for the entire page, the API will filter out any non-CORS third-party frames from the stack trace, keeping only the direct entry point called by the first-party script.

For example:

  1. myFirstPartyFunction()
  2. thirdPartyEntry() ⬅ Boundary kept (Actionable for the developer)
  3. [Filtered non-CORS frames] ⬅ Internal state hidden (Privacy maintained)

This surgical truncation ensures developers know exactly which external call blocked the thread, while completely hiding the internal execution state of that opaque script.

Stakeholder Feedback / Opposition

  • Chrome/WebUI : Positive (Driving initial incubation)
  • Web Performance WG : Invited for discussion

References & acknowledgements

Many thanks for valuable feedback and advice from:

  • Fergal Daly
  • Yoav Weiss

Appendix: Implementation Feasibility (Chromium PoC)

Demonstrating that continuous task monitoring and mid-execution interrupts can be achieved with near-zero overhead is critical to the viability of this proposal.

While the exact implementation details are left to each browser engine, we built a Proof of Concept (PoC) in Chromium to demonstrate that this interrupt-driven architecture is practically feasible.

The PoC utilizes a lock-free atomic timestamp on the main thread paired with a dedicated background watchdog thread. Our performance evaluation on a heavy workload (50,000+ micro-tasks and multiple >120ms synchronous tasks) yielded the following results:

  • Main-Thread Critical Path (< 50 ns): By using relaxed atomic stores (std::memory_order_relaxed) and reusing existing task scheduler timestamps, the observation overhead per task can be reduced to under 50 nanoseconds.
  • Watchdog Thread Overhead (~0.4% Duty Cycle): A background thread polling every 5ms takes approximately 20 µs per poll. This results in a background CPU utilization of just 0.4% when active.
  • Engine Interrupt Cost (< 6% Relative Overhead): Capturing the in-flight JavaScript engine stack trace takes a median of ~0.31 ms. Because this penalty is strictly deferred until a task has already violated the performance budget (e.g., 5ms), it represents less than 6% relative overhead on the janky task itself. Well-behaving tasks incur exactly 0 ms of capture overhead.

These results prove that capturing point-in-time stack traces via a watchdog thread is practically feasible and safe for continuous, in-the-wild production monitoring.

About

Explainer for the Task Interrupt Timing API: PerformanceObserver entry to detect main-thread jank with configurable thresholds and point-in-time stack trace.

Resources

Code of conduct

Contributing

Stars

1 star

Watchers

2 watching

Forks

Releases

Packages

Used by

Contributors

Languages