← TechDebtRoadmap · Blog

Tech debt radar: four axes as decision-support, not a certificate

A radar chart is easy to misuse. Stakeholders see a shape, assume the system is “fine” or “on fire,” and skip the evidence underneath. Used carefully, a four-axis tech debt radar is still one of the fastest ways to put code, architecture, test, and documentation health on the same slide. The chart’s job is prioritization clarity — not certification, not a promise of zero incidents, and not a substitute for reading scanner output.

Why four axes beat a single “debt score”

A single composite score hides trade-offs. A service can look acceptable on average while test debt on the checkout path is Critical. Splitting the picture into four axes forces the conversation onto which capability is actually blocking delivery. Code debt without architecture debt often means local cleanups will stick. Architecture debt without test debt often means refactors will land blind. Documentation debt rarely pages anyone at 2 a.m., yet it silently taxes every onboarding week.

Quality frameworks such as ISO/IEC 25010 remind teams that maintainability, reliability, and security are distinct characteristics. Your radar does not need to mirror every characteristic in the standard. It needs enough separation that two managers arguing about “the debt” are no longer talking past each other.

How to read current vs target polygons

Plot current scores as one polygon and optional targets as another. The gap between them is the capacity conversation for the quarter. Large gaps on test plus architecture usually dominate feature throughput more than a moderate code-debt gap on a quiet module. When target equals current on every axis, either the team is coasting or the targets are not ambitious — both are useful facts for leadership if stated honestly.

Keep the scale consistent: 0–100 with higher meaning healthier. Inverting a scale mid-year (so that “higher debt” reads as “higher score”) breaks trend comparisons and confuses anyone who joins mid-cycle.

Honesty boundaries you should print under the chart

These boundaries align with how Cunningham’s metaphor and Fowler’s quadrant were meant to be used: as thinking tools for trade-offs, not as marketing seals.

Pair the radar with actions that include rollback

A pretty polygon without an action table is decoration. Each selected action should state effort, expected benefit on a named axis, risk during change, and rollback. Rollback is where many debt programs fail — teams merge a “cleanup” that breaks a subtle workflow and have no path back. If rollback is “git revert” only, say so. If rollback needs a feature flag or dual-write period, write that down before the sprint starts.

Prompt injection and narrative risk

When a language model drafts the stakeholder narrative, treat the text as untrusted assistance. OWASP’s guidance on LLM risks (including prompt injection as LLM01) is a useful reminder that pasted audit notes can contain adversarial instructions. Prefer the deterministic table for capacity decisions; use prose as explanation, not as authority. See the OWASP Top 10 for LLM Applications.

Operating cadence

Re-score on a fixed cadence — monthly for fast-moving products, quarterly for slower platforms. After each incident, update the architecture and test axes even if the code axis looks unchanged. Publish the radar in the same place every time so trends are visible. When an axis stays Critical for two cycles, escalate the process issue (hiring, ownership, CI policy), not only another refactor ticket.

Indie teams can run the same loop with lighter ceremony: one Cursor audit pass, paste JSON, export the table into Linear, and revisit after the next launch. Enterprise teams may attach DPAs and private model keys; the honesty rules stay the same.

Design choices behind the four axes

Code and architecture are separated because cleaning methods inside a module rarely fixes illicit cross-module reads. Test and documentation are separated because green unit coverage can coexist with missing operator runbooks. Collapsing those pairs into two axes would recreate the “single score” problem at a slightly higher resolution. Four is a compromise: enough for QBR slides, few enough that people remember the labels without a legend tattoo.

Angles on the chart are fixed so trends remain comparable. If you reorder axes between quarters, historical screenshots lie. Keep the same order in every export and in every blog or wiki embed.

Facilitating the meeting around the radar

Open with the honesty line: decision-support, not a certificate. Ask each stakeholder to name the axis they believe is most expensive before you reveal scores. Divergence between gut feel and measured axes is itself a finding — it often means evidence is missing or incentives reward shipping over repayment. Then walk the gap to target, not the absolute height of every spoke.

Time-box debate on any single action to a few minutes. If rollback cannot be stated, park the item. If effort exceeds remaining capacity, split the action or defer it with a written reason. End by recording owners and re-score dates in the same document that holds the table.

Remote-friendly and async usage

Distributed teams can paste the radar PNG or SVG into an async update and link the action table. Comments should challenge evidence, not aesthetics. When AI narrative is enabled, paste it below the table with a clear Model-assisted label so readers know which sentences are explanatory prose. If the narrative disagrees with the table, trust the table until a human reconciles the difference.

For open-source maintainers, the same chart can communicate roadmap intent to contributors: “this quarter we raise test from 35 toward 65; drive-by feature PRs that increase architecture debt will be declined.” Clarity reduces contributor frustration more than another generic CONTRIBUTING paragraph.

What success looks like after two quarters

Success is not a perfect diamond on the radar. Success is a boring meeting where Critical axes have owners, deferred items have dates, and inventing metrics is socially unacceptable. Feature throughput should become more predictable because surprises move from production into the planning table. If scores rise while lead time worsens, your axes may be measuring vanity — revisit evidence sources rather than celebrating the polygon.

Keep the Cursor prompt and scanner versions noted beside each re-score so audits remain reproducible. When you change scoring rules, bump a RULESET date in your notes. Stale rules are a form of documentation debt on the governance process itself.

Accessibility of the radar for mixed audiences

Not every stakeholder reads charts fluently. Pair the polygon with a short table of axis, score, severity band, and top action. Colour alone should not carry meaning; include the numeric score in text. When exporting to PDF or wiki, keep the honesty disclaimer within one scroll of the chart so screenshots forwarded in chat still carry the decision-support caveat.

For executive audiences, lead with the business outcome tied to the worst axis. For senior engineers, lead with evidence sources and rollback. Same radar, different narrative emphasis — without changing the underlying numbers.

Comparing radar snapshots across services

Platform groups often maintain many services. Resist averaging radars into one fleet score without showing variance. A healthy average can hide one Critical payment service. Prefer a small gallery of radars with the same scale, or a table of worst-axis-per-service. When two services share a library, architecture debt in the library should appear on both charts or on an explicit shared-component row so ownership is not ambiguous.

Contract tests between services belong on the test axis of each consumer as well as the provider. Documentation debt for a shared OpenAPI file should list a single owner even if many teams consume it. Clear ownership prevents the radar from becoming a blame chart.

When to stop using the radar temporarily

During an active severity-one incident, pause roadmap theatre. Stabilize first, then update architecture and test axes with facts from the incident review. Using the radar as a weapon in a blame meeting destroys the psychological safety required for honest scores. Return to the cadence only when the goal is planning, not punishment.

Keep publishing the same honesty line under every chart export so forwarded screenshots still carry the decision-support caveat for new readers.

FAQ

Does a green radar mean the system is certified healthy?

No. A radar summarizes self-reported or scanner-backed scores for planning. It is not a compliance certificate.

Why include documentation as a debt axis?

Stale API docs and missing onboarding runbooks slow every new hire and increase change risk, even when code looks tidy.

Can AI write the narrative for the radar?

Yes as an optional labelled assist. If AI is unavailable, fail closed or stay on the deterministic table — never silent fake success.

Related: How it works · Quarterly roadmap guide · FAQ