Platform Engineering Metrics Great Teams Track Every Week

Platform Engineering Metrics Great Teams Track Every Week
Platform Engineering Metrics Great Teams Track Every Week

Platform teams are surrounded by things that are easy to count: portal visits, templates, clusters, pipelines, tickets, deployments, cloud spend, documentation pages, and the number of services registered in a catalog. Most of those numbers can move while the developer experience gets worse.

The useful question is not “What can the platform team measure?” It is “What signals tell us, every week, whether the platform is becoming easier to adopt, more successful to use, more reliable to depend on, and more effective at removing work?”

That distinction matters because an internal platform is not merely infrastructure. Current DORA platform-engineering guidance treats the platform as an internal product and recommends a balanced scorecard spanning software delivery performance, developer satisfaction, adoption and retention, and task success. The CNCF Platforms White Paper similarly argues for quantitative and qualitative measurement across user experience, organizational efficiency, and product delivery.

The result is not a single “platform KPI.” It is a small operating system of metrics.

Weekly metrics should answer operational questions

A metric belongs in a weekly platform review when a material change in the number could alter a decision.

That test immediately filters out much of the usual dashboard clutter.

If active teams using the deployment path fall sharply, the team should investigate adoption or reliability. If provisioning latency rises, the workflow may need engineering attention. If support tickets increase while “self-service adoption” rises, the platform is probably shifting work rather than eliminating it. If the platform exhausts an error budget, reliability work should displace feature work. If delivery outcomes deteriorate for teams using the platform, the team should look for a regression rather than celebrate portal traffic.

By contrast, a metric that is repeatedly displayed but never changes priorities is decoration.

This is why strong weekly scorecards are intentionally small. The goal is not observability of every internal component. It is decision-quality visibility into the platform as a product.

1. Active team adoption, not raw user traffic

The first signal is whether teams are actually using the platform capabilities that matter.

Raw daily active users or page views can help with instrumentation, but they are weak measures of value. A developer portal can accumulate traffic because engineers are repeatedly searching for information they cannot find. A CLI can execute frequently because a workflow requires five commands instead of one. High engagement can represent friction.

Backstage’s own adoption guidance is unusually explicit about this. It documents usage metrics and notes that Spotify sent weekly digests showing changes in plugin usage, while also warning against optimizing engagement for its own sake.

A better weekly view starts with teams and capabilities:

  • active teams using each critical platform capability;
  • new teams adopting it;
  • teams that stopped using it;
  • retained teams returning to use it;
  • share of eligible services or teams using the supported path.

The denominator is essential. “Forty teams used the deployment platform this week” means little if the addressable population is forty-five teams, and something very different if it is four hundred.

For capabilities used infrequently, use rolling windows. A database-provisioning workflow may be healthy even if no team runs it in a particular week. The scorecard can still be reviewed weekly while the underlying metric uses a longer window.

Instrument adoption at capability level

Avoid one global “platform user” flag. It hides which parts of the platform are working.

A team might use the catalog but bypass deployment automation. It may use the secret-management path but provision databases manually. Another team may have migrated completely to the golden path except for one regulated workload.

Represent adoption as a capability matrix:

Team Service template CI/CD Runtime Observability Database Secrets
Team A Active Active Active Active Active Active
Team B Active Active Active Active Manual Active
Team C Inactive Active Active Partial Manual Active

That matrix supports a better product conversation. “Team B has not adopted the platform” is too vague. “Team B still provisions databases outside the supported path because the platform does not cover its replication requirement” is actionable.

2. Task success for the journeys that justify the platform

Adoption without task success is not success.

DORA explicitly includes task success in its recommended platform scorecard. The principle also aligns with Google’s HEART framework, which connects product goals to user-centered behavioral measures.

For a platform team, the right unit is usually a developer journey rather than a portal page or API endpoint.

Examples include:

  • create a new service;
  • provision a database;
  • create a test environment;
  • deploy to production;
  • rotate a secret;
  • onboard an existing service to observability;
  • request an approved cloud capability;
  • diagnose a failed deployment.

For each critical journey, instrument the start, completion, failure, retry, cancellation, and escalation to a human. Then review a small set of weekly measures:

Task success rate. What proportion of started workflows finish successfully?

Time to success. How long does the workflow take at the median and at a tail percentile?

Retry rate. How often do developers need to repeat part of the workflow?

Abandonment rate. How often is a workflow started but not completed?

Escalation rate. How often does a supposedly self-service journey require a platform engineer?

These measures reveal problems that aggregate availability does not. A provisioning API can be technically “up” while users repeatedly fail policy validation with unclear errors. From the platform customer’s perspective, that is a broken product.

Measure the full journey, not the happy-path API

Platform instrumentation often stops too low in the stack.

Suppose the platform exposes a database-provisioning API. The API returns HTTP 200, the infrastructure controller accepts the request, and the metrics show perfect availability. But the developer still has to wait for an approval in another system, discover the generated credentials manually, and ask a platform engineer which network policy to apply.

The API succeeded. The task did not.

A task-success event should therefore represent the point at which the developer can actually use the capability. That may require correlating events across the portal, API, workflow engine, cloud provider, identity system, and deployment system.

Use a journey or correlation identifier from start to finish. Record state transitions, not just HTTP requests.

A minimal event model can contain:

  • journey type;
  • consuming team or anonymized team identifier;
  • service or workload identifier where appropriate;
  • start timestamp;
  • terminal timestamp;
  • terminal state;
  • failure category;
  • retry count;
  • manual-intervention flag;
  • platform version or workflow version.

Do not collect sensitive source code, secrets, or payload contents merely because the telemetry system makes it easy.

3. Manual-touch rate and exception demand

Self-service is often claimed too early.

The CNCF Platform Engineering Maturity Model distinguishes custom processes, standard tooling, self-service solutions, and integrated services. The key difference is not whether a portal or CLI exists. It is how much human intervention remains in the path.

That makes manual-touch rate one of the most useful weekly measures a platform team can create:

Of the workflows intended to be self-service, what proportion required someone from the platform team to intervene?

Pair it with exception demand:

  • new exception requests;
  • open exception backlog;
  • age of unresolved exceptions;
  • repeat exception categories;
  • share of requests that fall outside the supported path.

This metric family prevents a common false positive. Adoption can increase while the platform team becomes a larger operational bottleneck. In that situation the platform is spreading, but leverage is not.

The purpose of platform engineering is not to make a centralized team process more tickets through a nicer interface. It is to encode repeatable capability so application teams can move without waiting.

Exceptions are product discovery data

Do not treat every exception as a failure to enforce standards.

An exception can reveal:

  • a legitimate workload class the platform does not support;
  • an unnecessary platform constraint;
  • a missing parameter or escape hatch;
  • an organizational policy that should be automated;
  • poor documentation;
  • a product team bypassing the platform for no defensible reason.

Classify exceptions by cause and destination. Some should produce a platform feature. Some should produce a policy change. Some should remain exceptions because the use case is rare and costly to generalize.

Weekly exception data is valuable because it tells the platform team where its abstraction boundary is breaking.

4. Fulfillment latency for common platform capabilities

CNCF guidance specifically recommends measuring latency from a request to fulfillment of a service or capability and the time required to build and deploy a new service.

This is where product analytics meets queueing reality.

For each important workflow, capture end-to-end elapsed time rather than only internal execution time. If infrastructure provisioning takes minutes of compute but sits in an approval queue for much longer, the developer experiences the whole delay.

Review the distribution, not only the average:

  • median fulfillment time;
  • p95 or another tail measure suitable for your volume;
  • queue time versus execution time;
  • time spent waiting for approval;
  • time spent waiting for a platform engineer;
  • time lost to retries.

The objective is not to force every workflow toward zero seconds. Some actions legitimately require review. The metric exists to expose where time accumulates and whether that delay is intentional.

Separate platform latency from external latency

A useful decomposition is:

Platform processing time: time spent in platform-owned services.

Dependency time: cloud APIs, security scanners, artifact repositories, identity providers, or external systems.

Queue/approval time: waiting for a human or a policy gate.

User-recovery time: time between an error and a successful retry.

The developer experiences the sum. The engineering owner needs the components.

This lets the weekly review answer two different questions: “Is the developer journey getting slower?” and “Which part can the platform team change?”

5. Platform reliability and error-budget consumption

Internal platforms are production systems. Their customers happen to be other engineers.

That means reliability should be measured through user-visible service levels, not only Kubernetes pod health or API uptime. Google’s SLO guidance emphasizes service-level indicators and objectives as user-focused ways to reason about reliability, while error budgets provide a mechanism for balancing change against reliability work.

For a platform, useful SLIs might include:

  • successful production deployments through the platform;
  • successful environment creation within an expected time;
  • successful secret retrieval;
  • successful CI execution attributable to platform components;
  • successful catalog or dependency lookup;
  • successful authentication for developer-facing interfaces.

The weekly review should show SLO attainment and error-budget burn for the critical journeys.

This changes reliability discussions. Instead of asking whether the platform “had incidents,” the team asks whether developers could complete the work the platform promises to enable.

Do not build SLOs around components developers cannot see

An API availability SLO is useful when that API is the product boundary. It is less useful when the user journey spans many services and the API can be healthy while the outcome fails.

Choose SLIs close to user intent. “Deployment workflow reached production successfully within the expected window” is often more meaningful than “workflow-controller process returned successful health checks.”

Then use component metrics to diagnose the SLI when it degrades.

The SLO should be a product contract. The telemetry underneath is an engineering diagnostic system.

6. Support load and toil

A platform can appear healthy while consuming its own team.

Google SRE’s guidance on measuring toil is useful here: recurring operational burden should be measured objectively so teams can prioritize automation and determine whether toil-reduction work actually helped.

A platform team should therefore track:

  • support requests by category;
  • manual interventions;
  • pages and urgent operational interruptions;
  • recurring requests that should become product capabilities;
  • engineering time consumed by repetitive operational work.

Avoid converting this into invasive individual time tracking. The goal is to identify system-level demand.

The important pattern is the relationship between adoption and support. Healthy platform growth should eventually reduce manual work per consuming team. If adoption grows and support demand grows at the same rate, the platform has not yet created much leverage.

Normalize support demand

Absolute ticket counts are deceptive when the user population is changing.

Track ratios such as:

  • support requests per active team;
  • manual interventions per successful workflow;
  • urgent interruptions per platform capability;
  • exception requests per hundred eligible workflows, when volume supports the denominator.

The denominator turns growth into a comparable signal.

If total tickets rise because platform adoption doubled while tickets per active team fell, the platform may be scaling successfully. If both rise, the architecture or product boundary may be creating operational debt.

7. Software delivery outcomes

Platform engineering ultimately exists to improve the conditions under which teams deliver software.

DORA’s current software delivery performance metrics cover change lead time, deployment frequency, failed-deployment recovery time, change fail rate, and deployment rework rate. DORA’s 2026 history of the metrics explains how the model evolved to the current five measures.

These belong in the platform scorecard, but with an important warning: they are not proof that the platform caused the result.

Deployment frequency can rise because product teams changed branching practices. Lead time can improve because test suites became faster. Recovery time can change because observability improved independently. Organizational design, product architecture, team experience, and application maturity all affect delivery.

Use DORA metrics as outcome signals, then investigate.

Where the data permits it, segment by platform capability, migration cohort, or team adoption state. The question is not “Can we claim ROI from this chart?” It is “Did the intended outcome move in the expected direction, and what else changed at the same time?”

Avoid averaging away the teams that need help

A company-wide lead-time average can improve while several teams deteriorate.

Use distributions and cohorts:

  • teams fully on the supported path;
  • teams partially migrated;
  • teams outside the platform;
  • new adopters;
  • long-term adopters;
  • workload classes with different release characteristics.

Do not rank teams against one another without context. A regulated batch system, a mobile application, and a continuously deployed web service can have legitimately different delivery profiles.

The scorecard is a diagnostic tool, not a competition.

8. Developer sentiment, at the right cadence

Developer satisfaction matters, but weekly surveys usually do not.

DORA recommends regular developer satisfaction measures, and the CNCF white paper similarly includes user-satisfaction surveys. Frameworks such as HEART and SPACE exist partly because productivity cannot be reduced to activity counts.

The weekly platform scorecard can show the most recent satisfaction result without collecting a new survey every seven days.

For many organizations, monthly or quarterly sampling will produce better data. The exact cadence depends on population size, change frequency, and survey fatigue.

Useful questions are concrete:

  • Can you complete common workflows without asking another team?
  • Can you understand failures well enough to recover yourself?
  • Do platform defaults help or obstruct your work?
  • Can you find the capability or documentation you need?
  • Do you trust the supported path for production work?

Open-text responses are often more actionable than a single satisfaction score.

Pair sentiment with behavior

Sentiment and behavioral data can disagree, and that disagreement is useful.

Developers may report that a workflow is frustrating even while task-success rates are high. The issue may be cognitive load, poor feedback, or unnecessary steps.

Conversely, developers may like a polished portal while continuing to bypass it for production work. Satisfaction alone does not prove adoption or value.

The strongest picture comes from combining what developers say with what they can successfully do.

9. Cost and efficiency, when they are part of the platform promise

Cost belongs in the scorecard when the platform explicitly exists to improve infrastructure efficiency, provide standardized economic guardrails, or make cost visible to consuming teams.

But raw cloud spend is usually a poor weekly platform KPI.

Spend can rise because traffic, data volume, customer count, or model usage rises. Spend can fall because a product is shrinking. Neither direction establishes platform quality.

Use contextual measures such as:

  • cost per workload unit that the business already understands;
  • cost per environment or service class;
  • idle or stranded resource trends;
  • cost of platform-owned shared services;
  • cost of repeated exception paths;
  • unit-cost deltas after a platform migration, with workload changes documented.

The purpose is not to celebrate a lower bill independently of reliability and delivery. It is to understand whether the platform is helping teams consume infrastructure economically without reintroducing manual gates.

Build a metric contract before building the dashboard

Most measurement failures begin before the visualization layer.

For each metric, write a small contract:

Field Example
Name Deployment journey success rate
User journey Deploy application to production
Numerator Successful completed deployment journeys
Denominator All eligible deployment journeys reaching terminal state
Window Rolling four weeks, reviewed weekly
Segments Team, platform workflow version, workload class
Owner Platform delivery capability owner
Source Workflow events + deployment controller
Exclusions Explicit user cancellations before execution
Action Investigate when trend shifts materially or error budget burns rapidly

The contract prevents a silent definition change from becoming an apparent performance improvement.

It also makes instrumentation reviewable. Engineers can challenge whether the denominator is correct, whether retries are double-counted, whether automated retries should be visible, and whether a “success” event occurs too early.

Instrumentation architecture for a platform scorecard

The simplest useful architecture has four layers.

1. Event collection

Emit product events from the developer-facing journeys: portal, CLI, API, templates, CI/CD workflows, provisioning controllers, and service catalog.

Events should describe user intent and state transitions.

2. Operational telemetry

Collect metrics, logs, and traces for platform services. These diagnose why a user-facing task failed or slowed down.

3. Delivery telemetry

Collect software-delivery measures from source control, CI/CD, deployment systems, and incident systems. Keep definitions aligned with the metric you claim to report.

4. Curated metric layer

Transform raw events into stable metric definitions with versioned queries or data models. The weekly dashboard should read from this curated layer rather than embedding business logic in each panel.

That separation matters. Raw telemetry changes constantly. A metric used for management decisions needs a stable definition.

What not to optimize

Several easy metrics should remain diagnostic, not goals.

Number of platform features

More capabilities can increase cognitive load, maintenance cost, and fragmentation. A platform with fewer well-integrated paths may deliver more value.

Portal traffic

Traffic can signal adoption, confusion, or workflow complexity. Never optimize time spent in the portal.

Tickets closed

A platform team should want recurring ticket categories to disappear. High ticket throughput can reward the opposite behavior.

Infrastructure utilization in isolation

Efficiency matters, but a platform can reduce infrastructure utilization while making delivery slower or reliability worse. Cost metrics need workload and outcome context.

Deployment frequency alone

Increasing deployment frequency while change failure or rework rises is not an unambiguous improvement. DORA’s throughput and instability measures are designed to be considered together.

Story points, commits, and lines of code

These are activity measures, not platform outcomes. They are easy to game and difficult to compare across work types. A platform team should not justify itself by claiming it produced more code.

A practical weekly platform scorecard

A useful review can fit on one page.

Dimension Weekly signal Question
Adoption Active and retained teams by critical capability Are teams choosing the platform?
Task success Success, retry, abandonment, escalation Can developers finish the job?
Self-service Manual-touch rate and exception backlog Are we removing dependency on the platform team?
Flow Median/tail fulfillment latency Where is developer time waiting?
Reliability SLO attainment and error-budget burn Can teams trust the platform?
Support/toil Repeated requests and interventions Are we creating leverage?
Delivery Selected DORA trends Are downstream outcomes moving safely?
Sentiment Latest representative DevEx sample Do users experience the improvement?
Economics Contextual unit-cost or waste signal, if relevant Is the platform improving economic efficiency?

Each metric needs an owner and an action rule. If it moves materially, somebody should know what to inspect next.

Run the review as a control loop

The weekly platform review should not become a status meeting where every chart needs to be green.

A useful review asks five questions:

  1. What changed?
  2. Is the change real or sampling noise?
  3. Which user journey explains it?
  4. Does it require reliability work, product work, documentation, policy work, or removal of a capability?
  5. What decision follows from the evidence?

Keep the meeting centered on exceptions and movement. Stable metrics do not need narration.

A mature review should produce decisions such as:

  • pause rollout because the deployment workflow is burning its error budget;
  • prioritize a recurring exception because it now dominates manual interventions;
  • remove a low-adoption capability whose maintenance burden exceeds its value;
  • investigate a drop in retained teams after a workflow change;
  • split a slow provisioning journey into platform, external dependency, and approval components;
  • run targeted discovery with teams whose task-success rate fell;
  • defer a feature because support load shows the team needs automation first.

That is the difference between a dashboard and a management system.

There is no authoritative target saying every internal platform should have a particular adoption percentage, self-service rate, ticket volume, or provisioning latency. The correct baseline depends on the user population, risk profile, architecture, workload, and starting point.

Comparisons can still be useful, but internal trend lines are often more actionable than external benchmarks.

The deeper discipline is to resist false precision. Measure enough to detect change. Segment enough to understand it. Keep definitions stable enough to trust it. Combine product, reliability, operational, and delivery signals so one metric cannot dominate the story.

Great platform teams do not measure everything every week. They review a small set of signals that tell them whether the platform is earning adoption, completing developer tasks, eliminating human dependency, staying reliable, and improving the conditions for software delivery. Everything else is telemetry until it becomes a decision.

Also read: