# Uptime, latency, and SLOs

> For the complete documentation index, see [llms.txt](https://docs.temporal.io/llms.txt).
> Any documentation page is available as raw Markdown by appending `.md` to its URL.

> Temporal Cloud's 99.99% uptime objective, the 200ms p99 latency SLO, Visibility availability, and measured latency for starting and signaling Workflows.

> **📝 Note:**
>
> This page describes Temporal's current operational practices and does not create any additional commitment, representation, or warranty.
>

Temporal maintains multiple internal [service level objectives](https://en.wikipedia.org/wiki/Service-level_objective) (SLOs) and an internal alerting system based on those SLOs.
The alerting system accounts for all errors, not only the service errors that count against the [Temporal Cloud SLA](https://temporal.io/sla), and also monitors latency.
Temporal personnel receive an alert when an SLO is not being met, and on-call engineers are paged, which often means that issues are resolved before you notice them.

For current system status and recent incidents, see [Temporal Status](https://status.temporal.io).
For the regions you can run in, see [Service regions](/evaluate/cloud/regions); for the full set of rate, resource, and configuration limits, see [System limits](/evaluate/cloud/limits).

## Service levels at a glance 

| Measure                                                 | Applies to                                                                                                                                      | Target                                                | Type            |
| ------------------------------------------------------- | ----------------------------------------------------------------------------------------------------------------------------------------------- | ----------------------------------------------------- | --------------- |
| [Uptime](https://temporal.io/sla)                       | Covered APIs, per Namespace                                                                                                                     | See the [Temporal Cloud SLA](https://temporal.io/sla) | Contractual SLA |
| [Uptime](#uptime)                                       | All Namespaces                                                                                                                                  | 99.99%                                                | SLO             |
| [Latency](#latency)                                     | Worker requests, p99 per region                                                                                                                 | 200ms                                                 | SLO             |
| [Visibility API availability](#visibility-availability) | List and Count requests, rolling 30-day window                                                                                                  | 99.5%                                                 | SLO             |
| [Cloud Ops API availability](/ops#availability)         | HTTP and gRPC requests, monthly                                                                                                                 | 99.99%                                                | SLO             |
| [Recovery Time Objective (RTO)](/cloud/rpo-rto)         | Namespaces using High Availability with automatic failover, for the [outage types their replication covers](/cloud/rpo-rto#rto-and-rpo-summary) | Under 20 minutes                                      | Objective       |
| [Recovery Point Objective (RPO)](/cloud/rpo-rto)        | Namespaces using High Availability with automatic failover, for the [outage types their replication covers](/cloud/rpo-rto#rto-and-rpo-summary) | Under 1 minute                                        | Objective       |

## Uptime 

Temporal Cloud operates every Namespace to a 99.99% [service availability](https://en.wikipedia.org/wiki/Reliability,_availability_and_serviceability) objective, whether or not the Namespace uses [High Availability](/cloud/high-availability).

The contractual uptime commitment is the [Temporal Cloud Service Level Agreement (SLA)](https://temporal.io/sla).
It defines which APIs are covered, how availability is calculated, what is excluded, and the service credits that apply.
Namespaces using High Availability have a higher SLA than standard single-region Namespaces.

Temporal Cloud runs a cell architecture.
Each cell holds the software and services needed to host a Namespace, and the components in a cell are distributed across at least three availability zones per region.
For how Temporal Cloud handles each type of outage, see [Outages and Recovery Objectives](/cloud/rpo-rto).

## Latency 

Temporal Cloud has a p99 latency SLO of 200ms per region.

That SLO covers normal Worker requests, meaning commands and polling, and it applies to Nexus in both the caller and handler Namespaces.

### Measured latency 

Measured latency over a week-long period for starting and signaling Workflow Executions, and for starting [Standalone Activities](/standalone-activity) (`StartActivityExecution`),
was as follows:

#### August 2026

| Operation                          | p50  | p90  |   p99 |
| :--------------------------------- | :--: | :--: | ----: |
| `StartWorkflowExecution`           | 20ms | 32ms |  78ms |
| `SignalWorkflowExecution`          | 19ms | 42ms |  91ms |
| `SignalWithStartWorkflowExecution` | 30ms | 47ms | 109ms |
| `StartActivityExecution`           | 13ms | 19ms |  45ms |

#### Earlier measurements

**January 2026**

| Operation                          | p50  | p90  |  p99 |
| :--------------------------------- | :--: | :--: | ---: |
| `StartWorkflowExecution`           | 14ms | 21ms | 69ms |
| `SignalWorkflowExecution`          | 11ms | 19ms | 46ms |
| `SignalWithStartWorkflowExecution` | 19ms | 37ms | 95ms |

**March 2024**

| Operation                          | p90  |  p99 |
| :--------------------------------- | :--: | ---: |
| `StartWorkflowExecution`           | 24ms | 54ms |
| `SignalWorkflowExecution`          | 14ms | 40ms |
| `SignalWithStartWorkflowExecution` | 24ms | 61ms |

Latency observed from the Temporal Client also reflects other components in the path, such as the Codec Server, an egress proxy, and the network itself.
Concurrent operations on the same Workflow Execution can raise latency as well.

### Custom persistence layer

Temporal Cloud runs a custom persistence layer rather than the persistence stores available to self-hosted deployments.
Three components of that layer account for most of the difference in latency:

- **Sharding:** Distributes load across multiple databases and resizes them independently, so a traffic spike in one shard doesn't become a bottleneck for the rest.
- **Write-ahead log (WAL):** Batches updates in an append-only log before writing them to the database, which reduces write latency and database size.
- **Tiered storage of Event History:** Moves the Event History of closed Workflow Executions to cheaper storage, which keeps the primary database smaller and faster for running Executions.

## Visibility API availability 

Temporal Cloud targets 99.5% availability for Visibility read requests, measured over a rolling 30-day window.
This covers List and Count operations against Workflow Executions, Activity Executions, Nexus Operations, Schedules, and Batch Operations.

This target is deliberately lower than the availability guaranteed for core Workflow Service APIs.
Visibility is a search index intended for operational discovery, not a critical path for application logic.

This is a service level objective rather than a contractual commitment.
Visibility API requests are not covered by the [Temporal Cloud SLA](https://temporal.io/sla), and no service credits are associated with this objective.

The objective measures whether a request returns a valid response, not how current the results are.
For what it excludes and how to interpret it alongside Visibility's eventual consistency, see [Visibility availability in Temporal Cloud](/visibility#availability).

## Throughput 

Each Namespace has a rate limit measured in [Actions](/cloud/pricing#action) per second, and two capacity modes control it:

- **On-Demand Capacity** raises the Namespace limit automatically as usage grows.
- **Provisioned Capacity** sets the limit from the Temporal Resource Units (TRUs) you request, which suits predictable spikes and guaranteed throughput.

See [Capacity Modes](/cloud/capacity-modes) for how each mode sets your limit.

## Related

- [Temporal Cloud SLA](https://temporal.io/sla): The contractual uptime commitment, including covered APIs, calculation, exclusions, and service credits.
- [Outages and Recovery Objectives (RTO / RPO)](/cloud/rpo-rto): Recovery targets for each type of outage.
- [High Availability](/cloud/high-availability): Replication and failover options for a higher contractual SLA.
- [Higher throughput and lower latency: Temporal Cloud's custom persistence layer](https://temporal.io/blog/higher-throughput-and-lower-latency-temporal-clouds-custom-persistence-layer): How sharding, the write-ahead log, and tiered storage are built.
- [Benchmarking latency: Temporal Cloud versus self-hosted Temporal](https://temporal.io/blog/benchmarking-latency-temporal-cloud-vs-self-hosted-temporal): Measured comparison of the two deployment options.
- [Replay conference talk: Custom persistence layer](https://www.youtube.com/watch?v=SQv9ot-jB6o&list=PLl9kRkvFJrlREHL7fiEKBWTp5QuFeYS2r&index=5): Walkthrough of the persistence design.
