OSS Temporal Service metrics reference
The information on this page is relevant to open source Temporal Service deployments.
See Cloud metrics for metrics emitted by Temporal Cloud.
See SDK metrics for metrics emitted by the SDKs.
Temporal SDKs and the Temporal Service attach these metric dimensions as tags. Once a metric is scraped by Prometheus or another OpenMetrics-compatible system, the same dimensions are called labels.
A Temporal Service emits metrics that operators use to monitor the Temporal Service's performance and to set up alerts.
All metrics emitted by the Temporal Service are listed in metric_defs.go.
For details on setting up metrics in your Temporal Service configuration, see the Temporal Service configuration reference.
The dashboards repository contains community-driven Grafana dashboard templates that you can use as a starting point for monitoring the Temporal Service and SDK metrics. For any metrics that are missing in the dashboards, use metric_defs.go as a reference.
In addition to the metrics the Temporal Service emits, monitor infrastructure metrics such as CPU, memory, and network for every host that runs a Temporal Service service.
Each metric on this page lists its type:
- Counter: A value that only increases. Query it with
rate()orincrease(). - Gauge: A sampled value that can go up or down.
- Histogram: A distribution of values. Latency histograms are defined as timers in
metric_defs.go. Query them through the_bucketseries withhistogram_quantile().
All example queries on this page use PromQL.
Common metrics
Temporal emits metrics for each gRPC service request.
These metrics are emitted with type, operation, and namespace tags, which show request rates across services, Namespaces, and operations.
- Use the
operationtag in your query to get request rates, error rates, or latencies per operation. - Use the
service_nametag with a service role value, such asfrontend,history, ormatching, to get details for a specific service. The role values are defined in metric_defs.go.
All common tags that you can add in your query are defined in metric_defs.go.
For example, to see service requests by operation on the Frontend Service, use the following query:
sum by (operation) (rate(service_requests{service_name="frontend"}[2m]))
Start with the following metrics.
service_requests
Counts the gRPC requests the service receives, tagged by operation.
Type: Counter
Example: Request rate for the AddWorkflowTask operation
sum(rate(service_requests{operation="AddWorkflowTask"}[2m]))
service_latency
Shows latencies for all Client request operations. Start here to find which operation has high latency.
Type: Histogram
Example: P95 service latency by operation for the Frontend Service
histogram_quantile(0.95, sum(rate(service_latency_bucket{service_name="frontend"}[5m])) by (operation, le))
service_error_with_type
Counts errors the service encounters, tagged by error_type.
Type: Counter
Example: Service errors by type for the Frontend Service
sum(rate(service_error_with_type{service_name="frontend"}[5m])) by (error_type)
client_errors
Counts failed requests between Temporal Service roles. An increase indicates connection issues between roles.
Type: Counter
Example: Errors on requests from the Frontend Service to the History Service
sum(rate(client_errors{service_name="frontend",service_role="history"}[5m]))
The following sections list metrics for each service. For metrics not listed here, see metric_defs.go.
Matching Service metrics
poll_success
Counts Tasks that are successfully matched to a poller.
Type: Counter
Example:
sum(rate(poll_success{}[5m]))
poll_timeouts
Counts polls that return because no Task was available within the poll timeout.
Type: Counter
Example:
sum(rate(poll_timeouts{}[5m]))
asyncmatch_latency
Measures the time from creation to delivery for async matched Tasks. The larger this latency, the longer Tasks wait in the queue for your Workers to pick them up.
Type: Histogram
Example:
histogram_quantile(0.95, sum(rate(asyncmatch_latency_bucket{service_name="matching"}[5m])) by (operation, le))
no_poller_tasks
Counts Tasks added to a Task Queue that has no recent poller. An increase usually means that the Worker or the starter program is using the wrong Task Queue name.
Type: Counter
History Service metrics
A History Task is an internal Task that the Temporal Service creates as part of a transaction that updates Workflow state. The History Service processes these Tasks. Use the following metrics to monitor the health of History Task processing.
task_requests
Counts every Task process request.
Type: Counter
Example:
sum(rate(task_requests{operation=~"TransferActive.*"}[1m]))
task_errors
Counts every Task process error.
Type: Counter
Example:
sum(rate(task_errors{operation=~"TransferActive.*"}[1m]))
task_attempt
Records the number of attempts for each Task Execution. A Task is retried until it succeeds, and each retry increases the attempt count.
Type: Histogram
Example:
histogram_quantile(0.95, sum(rate(task_attempt_bucket{operation=~"TransferActive.*"}[1m])) by (operation, le))
task_latency_processing
Shows the processing latency per attempt.
Type: Histogram
Example:
histogram_quantile(0.95, sum(rate(task_latency_processing_bucket{operation=~"TransferActive.*",service_name="history"}[1m])) by (operation, le))
task_latency
Measures the in-memory latency across multiple attempts.
Type: Histogram
task_latency_queue
Measures the end-to-end duration from when the Task should run (when it fired) to when the Task is done.
Type: Histogram
task_latency_load
Measures the duration from Task generation to Task loading. This is the schedule-to-start latency for the persistence queue.
Type: Histogram
task_latency_schedule
Measures the duration from Task submission to the Task scheduler to processing. This is the schedule-to-start latency for the in-memory queue.
Type: Histogram
queue_latency_schedule
Measures the time to schedule 100 Tasks in one Task channel in the host-level Task scheduler. If fewer than 100 Tasks are in the Task channel for 30 seconds, the latency is scaled to 100 Tasks when emitted.
Type: Histogram
service_latency_userlatency
Shows the latency introduced by Workflow logic.
For example, one Workflow that schedules many Activities or Child Workflows at the same time can cause per-Workflow lock contention.
The time spent waiting for the per-Workflow lock counts as userlatency.
Type: Histogram
The operation tag on History Service metrics contains the Task type and whether the Task is Active or Standby.
Use it to get request rates, error rates, or latencies per operation, which helps identify issues caused by database problems.
Persistence metrics
The Temporal Service emits metrics for every persistence database read and write. Start with the following metrics.
persistence_requests
Counts every persistence request.
Type: Counter
Example: Persistence requests by operation for the History Service
sum by (operation) (rate(persistence_requests{service_name="history"}[1m]))
Example: Persistence requests by operation for the Matching Service
sum by (operation) (rate(persistence_requests{service_name="matching"}[1m]))
persistence_errors
Counts all persistence errors. An increase indicates connection issues between the Temporal Service and the persistence store.
Type: Counter
Example: Persistence errors for the History Service
sum(rate(persistence_errors{service_name="history"}[1m]))
persistence_error_with_type
Counts persistence store errors, tagged by error_type.
Type: Counter
Example: Persistence errors by type for the History Service
sum(rate(persistence_error_with_type{service_name="history"}[1m])) by (error_type)
persistence_latency
Shows the latency of persistence operations.
Type: Histogram
Example: P95 persistence latency by operation for the History Service
histogram_quantile(0.95, sum(rate(persistence_latency_bucket{service_name="history"}[1m])) by (operation, le))
Schedule metrics
The following metrics track the performance and outcomes of the Workflow Executions that Schedules start.
Schedule metrics are tagged by namespace, not by individual Schedule.
A non-zero value tells you that at least one Schedule in the Namespace is affected, but not which one.
To find the affected Schedule, see Troubleshoot missed Schedule Actions.
schedule_buffer_overruns
Counts the times the buffer that holds Scheduled Workflow Executions exceeds its maximum capacity.
This usually happens when a Schedule with the BufferAll overlap policy has an average run length longer than its average interval.
Type: Counter
Example:
sum(rate(schedule_buffer_overruns{namespace="$namespace"}[5m]))
schedule_missed_catchup_window
Counts the Scheduled Actions that the Temporal Service didn't run within the configured catchup window. An outage that lasts longer than the catchup window can cause missed Actions.
Type: Counter
Example:
sum(rate(schedule_missed_catchup_window{namespace="$namespace"}[5m]))
schedule_rate_limited
Counts the Workflow Executions a Schedule tried to start that were throttled by Namespace rate limits. Frequent rate limiting can cause missed catchup windows.
Type: Counter
Example:
sum(rate(schedule_rate_limited{namespace="$namespace"}[5m]))
schedule_action_success
Counts the Workflow Executions that a Schedule started successfully, on its schedule or from a manual trigger.
Type: Counter
Example:
sum(rate(schedule_action_success{namespace="$namespace"}[5m]))
Workflow metrics
These metrics count Workflow Execution outcomes.
workflow_cancel
Counts Workflow Executions canceled before completing.
Type: Counter
workflow_continued_as_new
Counts Workflow Executions that Continued-As-New.
Type: Counter
workflow_failed
Counts Workflow Executions that failed.
Type: Counter
workflow_success
Counts Workflow Executions that completed successfully.
Type: Counter
workflow_terminate
Counts Workflow Executions that were terminated.
Type: Counter
workflow_timeout
Counts Workflow Executions that timed out before completing.
Type: Counter
Nexus metrics
These metrics cover Nexus Operations.
Nexus machinery in the History Service
See the Nexus architecture document for how these components fit together.
In-memory buffer
dynamic_worker_pool_scheduler_enqueued_tasks(Counter): Increments when a Task is added to the buffer.dynamic_worker_pool_scheduler_dequeued_tasks(Counter): Increments when a Task is removed from the buffer.dynamic_worker_pool_scheduler_rejected_tasks(Counter): Increments when the buffer is full and a Task is rejected.dynamic_worker_pool_scheduler_buffer_size(Gauge): The sampled size of the buffer.
Concurrency limiter
dynamic_worker_pool_scheduler_active_workers(Gauge): The sampled number of running goroutines.
Rate limiter
rate_limited_task_runnable_wait_time(Histogram): The time a Task waits for the rate limiter.
Circuit breaker
circuit_breaker_executable_blocked(Counter): Increments each time the circuit breaker blocks a Task execution.
Task executors
nexus_outbound_requests(Counter): The number of outbound Nexus requests the History Service makes.nexus_outbound_latency(Histogram): The latency of outbound Nexus requests the History Service makes.callback_outbound_requests(Counter): The number of outbound callback requests the History Service makes.callback_outbound_latency(Histogram): The latency of outbound callback requests the History Service makes.
Nexus machinery on the Frontend Service
nexus_requests
The number of Nexus requests received by the service.
Type: Counter
nexus_latency
Latency of Nexus requests.
Type: Histogram
nexus_request_preprocess_errors
The number of Nexus requests for which pre-processing failed.
Type: Counter
nexus_completion_requests
The number of Nexus completion (callback) requests received by the service.
Type: Counter
nexus_completion_latency
Latency of Nexus completion (callback) requests.
Type: Histogram
nexus_completion_request_preprocess_errors
The number of Nexus completion requests for which pre-processing failed.
Type: Counter