Skip to main content

Uptime, latency, and SLOs

View Markdown
info

This page describes Temporal's current operational practices and does not create any additional commitment, representation, or warranty.

Temporal maintains multiple internal service level objectives (SLOs) and an internal alerting system based on those SLOs. The alerting system accounts for all errors, not only the service errors that count against the Temporal Cloud SLA, and also monitors latency. Temporal personnel receive an alert when an SLO is not being met, and on-call engineers are paged, which often means that issues are resolved before you notice them.

For current system status and recent incidents, see Temporal Status. For the regions you can run in, see Service regions; for the full set of rate, resource, and configuration limits, see System limits.

Service levels at a glance​

MeasureApplies toTargetType
UptimeCovered APIs, per NamespaceSee the Temporal Cloud SLAContractual SLA
UptimeAll Namespaces99.99%SLO
LatencyWorker requests, p99 per region200msSLO
Visibility API availabilityList and Count requests, rolling 30-day window99.5%SLO
Cloud Ops API availabilityHTTP and gRPC requests, monthly99.99%SLO
Recovery Time Objective (RTO)Namespaces using High Availability with automatic failover, for the outage types their replication coversUnder 20 minutesObjective
Recovery Point Objective (RPO)Namespaces using High Availability with automatic failover, for the outage types their replication coversUnder 1 minuteObjective

Uptime​

Temporal Cloud operates every Namespace to a 99.99% service availability objective, whether or not the Namespace uses High Availability.

The contractual uptime commitment is the Temporal Cloud Service Level Agreement (SLA). It defines which APIs are covered, how availability is calculated, what is excluded, and the service credits that apply. Namespaces using High Availability have a higher SLA than standard single-region Namespaces.

Temporal Cloud runs a cell architecture. Each cell holds the software and services needed to host a Namespace, and the components in a cell are distributed across at least three availability zones per region. For how Temporal Cloud handles each type of outage, see Outages and Recovery Objectives.

Latency​

Temporal Cloud has a p99 latency SLO of 200ms per region.

That SLO covers normal Worker requests, meaning commands and polling, and it applies to Nexus in both the caller and handler Namespaces.

Measured latency​

Measured latency over a week-long period for starting and signaling Workflow Executions, and for starting Standalone Activities (StartActivityExecution), was as follows:

August 2026​

Operationp50p90p99
StartWorkflowExecution20ms32ms78ms
SignalWorkflowExecution19ms42ms91ms
SignalWithStartWorkflowExecution30ms47ms109ms
StartActivityExecution13ms19ms45ms
Earlier measurements

January 2026

Operationp50p90p99
StartWorkflowExecution14ms21ms69ms
SignalWorkflowExecution11ms19ms46ms
SignalWithStartWorkflowExecution19ms37ms95ms

March 2024

Operationp90p99
StartWorkflowExecution24ms54ms
SignalWorkflowExecution14ms40ms
SignalWithStartWorkflowExecution24ms61ms

Latency observed from the Temporal Client also reflects other components in the path, such as the Codec Server, an egress proxy, and the network itself. Concurrent operations on the same Workflow Execution can raise latency as well.

Custom persistence layer​

Temporal Cloud runs a custom persistence layer rather than the persistence stores available to self-hosted deployments. Three components of that layer account for most of the difference in latency:

  • Sharding: Distributes load across multiple databases and resizes them independently, so a traffic spike in one shard doesn't become a bottleneck for the rest.
  • Write-ahead log (WAL): Batches updates in an append-only log before writing them to the database, which reduces write latency and database size.
  • Tiered storage of Event History: Moves the Event History of closed Workflow Executions to cheaper storage, which keeps the primary database smaller and faster for running Executions.

Visibility API availability​

Temporal Cloud targets 99.5% availability for Visibility read requests, measured over a rolling 30-day window. This covers List and Count operations against Workflow Executions, Activity Executions, Nexus Operations, Schedules, and Batch Operations.

This target is deliberately lower than the availability guaranteed for core Workflow Service APIs. Visibility is a search index intended for operational discovery, not a critical path for application logic.

This is a service level objective rather than a contractual commitment. Visibility API requests are not covered by the Temporal Cloud SLA, and no service credits are associated with this objective.

The objective measures whether a request returns a valid response, not how current the results are. For what it excludes and how to interpret it alongside Visibility's eventual consistency, see Visibility availability in Temporal Cloud.

Throughput​

Each Namespace has a rate limit measured in Actions per second, and two capacity modes control it:

  • On-Demand Capacity raises the Namespace limit automatically as usage grows.
  • Provisioned Capacity sets the limit from the Temporal Resource Units (TRUs) you request, which suits predictable spikes and guaranteed throughput.

See Capacity Modes for how each mode sets your limit.