Celestial Mode: Night

MFid Methodology v2.6.1 · 2026-08-05

One canonical formula. Two reporting forms. Published rubrics for every dimension.

For a worked example of the formula applied end-to-end on a published vendor SLA, see the MFid walkthrough. Prefer no formulas at all? The same ideas are explained without any math on the Plain English page.

Why this page exists

A firm whose entire product is auditing other people’s claims cannot ship a metric in two formulas, two evidence-tier names, and an undocumented Integrity rubric. This page is the canonical reference. Every other page on the site is required to match it. If a page contradicts this one, this one wins and the other page is a finding against ourselves.

This is the firewall recommended in the 2026-05-26 thesis review (§4.1, §4.2, §4.5). It exists because the review was right.

The canonical formula

MFidaggregate = (D × E × O × I)1/4

D, E, O, I ∈ [0,1]. Geometric mean of four normalized dimension scores.

Why a geometric mean and not an arithmetic mean. A geometric mean is dominated by the smallest factor. A single weak dimension caps the composite — an MFid of 0.5 on any one dimension limits the aggregate to roughly 0.84 even if the other three are perfect. That is the policy we want. An arithmetic mean would let a perfect Observability score paper over a broken Dependability score. We treat fidelity gaps as non-substitutable: you cannot fix a lie about latency by being more transparent about it.

How the dimensions interact under this aggregation. Equal relative leverage (a 1% relative move in any dimension moves the aggregate by ¼%), unequal absolute leverage (the weaker a dimension already is, the more each point of improvement is worth: ∂MFid/∂X = MFid / 4X), a hard ceiling set by the weakest factor (a dimension at x caps the aggregate at x1/4), and non-compensation near the top of the scale. The full derivation, with worked numbers, is on the Theory page § How the four dimensions interact.

Why these four dimensions and not three or five. D, E, O, I are the smallest set that survives every domain we have scored. Dependability, Efficiency, and Observability are the engineering core. Integrity is the dimension that distinguishes a fast wrong answer from a fast right one — the one most measurement frameworks omit and the one most needed in the autonomous-systems era.

Two reporting forms — and how they relate

Past versions of this site published two formulas without saying so. That was a finding. The reconciliation:

Form 1 — Aggregate MFid (the canonical number)

MFidaggregate = (D × E × O × I)1/4. One number per system, vendor, process, or stack. This is the number that goes on the board deck. It is always a geometric mean of four [0,1] dimension scores, and is therefore always in [0,1] itself — as of v2.4.0 that bound is a hard invariant of the instrument, with no exceptions. Weights are fixed at 1/4 each. There are no per-engagement weight choices.

Form 2 — Domain-projected fidelity (MFidapp, MFidnet, …)

When a system exposes domain-specific service-level indicators (latency, throughput, reliability for an app; bandwidth, jitter, loss for a network), we compute a domain projection:

MFidapp = LwL × TwT × RwR

where L, T, R ∈ [0,1] are fidelity scores per SLI (claimed/observed clipped to 1.0) and weights w sum to 1. The projection is a weighted geometric mean — the same non-compensatory operator as the aggregate. Through v2.3.0 this form was an arithmetic weighted sum, which quietly permitted inside a dimension the compensation the aggregate forbids between dimensions: a strong throughput score could offset a latency shortfall. That inconsistency was corrected in v2.4.0; see the Revision Ledger. Weights are not chosen per case. They are derived from the system’s own published SLO portfolio — if the operator weights latency at 50% of their SLO budget, we use 0.5. If no SLO portfolio exists, we use the default Tier-1 weighting (0.5 / 0.3 / 0.2 in order of business impact: response time, throughput, reliability) and publish the choice in the finding.

A domain projection is not the aggregate MFid. It is one input into one of the four dimensions. Roll-up:

  • Dependability (D) ← the per-SLI D term on each L/T/R signal — dispersion plus unfavorable offset against the claim, tail-capped; aggregate-type claims enter as raw fidelity ratios (see rubric below).
  • Efficiency (E) ← resource-bounded form of L and T (claimed cost-per-unit ÷ observed cost-per-unit).
  • Observability (O) ← coverage of L/T/R signals: fraction of customer journeys for which a measurement exists.
  • Integrity (I) ← scored separately; see rubric below.

Every published MFidapp in our case studies is annotated with both forms going forward: the projection that motivated the finding, and the aggregate it rolled into.

Terminology — the Software Defined Index (SDI)

SDI(x) ≡ MFid(x)

The Software Defined Index is the buyer-facing name for the canonical aggregate, answering the question the firm’s name poses: how software defined is X? — to what degree does X’s stated definition determine its behavior? SDI is not a third reporting form: it introduces no new aggregation, dimensions, weights, or evidence tiers. The two terms are equal by definition. If SDI ever diverges from MFid, the divergence must appear on the Revision Ledger before any diverged score is published. The derivation of the equality is on the Theory page; the naming argument is on Why “Software Defined”?.

Operational definitions of D, E, O, I

Each dimension is a normalized composite of underlying measurements. Each has a published formula. None of them are subjective. None of them are scored by vibes.

D — Dependability

Definition: Repeated observations under stated conditions match the claimed value within its stated tolerance — the buyer can depend on any single future observation.

Measurement: D = 1 − min(1, (Δ + σ) / (μc × τ)) where μc is the claimed value of the observable, σ is the standard deviation of the observations, τ is the published tolerance band (e.g. 10%), and Δ is the unfavorable offset of the observed mean from μc — zero when the observed mean meets or beats the claim, the size of the shortfall when it does not. Computed per SLI; aggregated by minimum across SLIs (a system is only as dependable as its worst-behaved indicator). For a claim that binds on aggregate rather than per observation — an uptime percentage, a durability figure — the SLI enters the minimum as its raw fidelity ratio (observed/claimed, clipped to 1): a 30-day window yields one observed uptime, not a distribution of them. The dispersion-plus-offset form applies to claims that bind on every observation.

Why the offset term (new in v2.4.0): dispersion-only scoring rewarded a system that is perfectly consistent at the wrong value — rock-steady at 200 ms against a 50 ms claim. Consistency at the wrong value is not dependability with respect to the claim; it is reliable unfaithfulness, and Δ makes it cost. Whenever the claim is met on average, Δ = 0 and the form reduces exactly to the prior σ / (μ × τ).

Worst-case binding: D is additionally bounded by a tail cap computed from the single worst observation in the window: D ≤ min(1, 2/z), where z is the largest excursion of any critical-SLI observation from the claimed value, in units of σ. The cap is inactive for z ≤ 2, passes the historical 0.9 near z ≈ 2.2, and keeps falling as the worst observation worsens — a 4σ excursion bounds D at 0.5. We refuse to let a calm hour hide a panic minute; and as of v2.5.0 we also refuse to score a 2.01σ event a cliff apart from a 1.99σ event — the penalty is continuous in the size of the excursion, and, unlike the old flat 0.9 cap, it does not stop charging at catastrophic ones.

E — Efficiency

Definition: Resource cost per unit of useful output, relative to the published or contracted cost. E is cost fidelity, not cost goodness: it asks “does it cost what they said?”, so an honestly-priced expensive system scores 1.0.

Measurement: E = min(1, claimed_cost_per_unit / observed_cost_per_unit), computed in the natural unit of the system (cycles/token, watts/inference, dollars/transaction, joules/request, bytes/query). One unit per system, declared up front. E is a one-sided ratio: under-cost (better than claim) is clipped at 1.0. Over-delivery is recorded in the finding — and, where it is structural, earns the Verified Underreporter annotation — but it never raises a score above 1.0.

O — Observability

Definition: Fraction of the customer-relevant behavior surface for which a current, queryable, retained measurement exists.

Measurement: O = (covered_SLIs / required_SLIs) × retention_factor × freshness_factor. Required SLIs are enumerated up front for each system from its specification (not from what is currently instrumented — that would let absence become a credit). Retention factor is 1.0 if telemetry is retained ≥ 30 days, scaled down otherwise. Freshness factor is 1.0 if the dashboard is queryable in < 60 seconds, scaled down otherwise. A score you have to mine from log files is not observability; it is archaeology.

Why O multiplies into the score rather than standing beside it as a confidence interval: a claim that cannot be observed is operationally indistinguishable from a claim that is false, so missing telemetry costs the vendor the same way a missed target does. The orthodox alternative — a headline score with an O-driven uncertainty band — carries the same information but lets an unmeasured system keep a high headline number. We refuse that trade, and the refusal is a design choice this page owns explicitly.

I — Integrity

This is the dimension the 2026-05-26 review flagged as undefined. The review was correct. The rubric below is the answer.

Definition: The fraction of system activity that demonstrably serves the stated purpose under audit, with the remainder classified as out-of-spec drift (not necessarily harmful — but not what was contracted).

Important framing. Integrity is not a property of the artifact in isolation; it is a property of the artifact relative to its specification. We measure the specification as carefully as we measure the system. An unclear spec produces a low Integrity ceiling, not a low Integrity score — we publish the ceiling and recommend the spec be tightened.

Measurement (three-part, each scored [0,1], aggregated by geometric mean):

  1. Ispec — Specification clarity. Does a written, dated, signed specification exist that enumerates required behaviors and forbidden behaviors? Scored by document analysis on a 7-point checklist (existence, dating, scope, behavior list, forbidden-list, change log, sign-off). Reproducible across reviewers with κ ≥ 0.7 on a 50-spec calibration set; calibration set published on request.
  2. Itrace — Operational coverage. Of the operations the system performed during the measurement window, what fraction can be traced to a specified behavior? Computed as (traced_operations / total_operations) from logs, traces, or transaction records. Operations with no trace match are not assumed malicious — they are assumed unscored, and counted against Itrace.
  3. Idrift — Forbidden-behavior detection. Of operations classifiable as “outside spec” (output schemas, data destinations, decision boundaries the spec excludes), what fraction were caught by automated guardrails before having effect? Computed as (blocked_violations / detected_violations). A system with no violations and no detection capability scores 0.5 (we cannot tell whether it is well-behaved or unmonitored).

I = (Ispec × Itrace × Idrift)1/3.

For ML and autonomous systems specifically: Ispec is scored against the model card and policy document. Itrace is computed from prompt/response logs against the policy classifier. Idrift measures jailbreak/red-team catch rate. No public worked example on an LLM endpoint has been published yet; until one is, the rubric above is the complete canonical definition.

What I is not. I is not a measure of whether the system is good. A perfectly-malicious system with a clear spec authorizing maliciousness scores I = 1.0. I measures fidelity to the spec, not the wisdom of the spec. That is by design. We score the gap between claim and reality; we do not score the claim itself. That is the customer’s job.

The honest underreporter — and why the score still stops at 1.0

The MFid ceiling is 1.0. Every dimension formula contains a min(1, …) clip, and a geometric mean of values bounded to [0,1] cannot exceed 1.0. As of v2.4.0 that ceiling is a hard invariant: no published MFid — and therefore no Software Defined Index — ever exceeds 1.0, under any conditions. Its purpose is twofold: a vendor that out-performs in one dimension must not mask a deficit in another, and an index that is bounded except when it isn’t is two definitions under one name — the exact defect this firm files against vendors.

There is a narrow category of entity for which the clip genuinely discards information: the verified, systematic underreporter — an organization whose published specifications are demonstrably and consistently conservative across multiple independent dimensions, not by accident, not once, but as a documented operating posture. Versions v2.2.0 through v2.3.0 of this methodology resolved that by lifting the clip and permitting published scores above 1.0. An external mathematical review (2026-08-04) found this to be the wrong resolution, and v2.4.0 withdraws it. The information the clip discards is real; it now lives in the annotation, not the score.

The canonical case: Porsche

Porsche publishes 0–60 times, quarter-mile figures, lap times, and power outputs that independent reviewers and owners consistently beat — not by a rounding error, but by a margin that is too regular to be noise. A 911 GT3 RS quoted at a Nürburgring time the factory already knows it can beat. A Taycan Turbo S rated at a 0–60 the launch-control software was tuned to exceed. The understatement is not accidental; it is policy. Porsche does not want a customer to discover their car underperforms the sticker. The sticker is a floor, not a target.

Applied to MFid: an audit of Porsche's published performance specifications against independently measured reality would find raw claimed-vs-observed ratios above 1.0 on every performance figure — the ratios that feed E and the per-SLI fidelity terms — before the clip is applied. D and I cannot exceed 1.0 by construction (one is 1 minus a clipped penalty, the other a geometric mean of fractions); they register the understatement differently, as near-perfect scores: the gap between claim and observation is consistent (high D), and the conservatism is documented policy rather than accident (high I). Under a bare clip the composite scores 1.0 — numerically indistinguishable from a vendor who just barely met their word. That conflation is the problem the annotation below solves.

The annotation rule for verified underreporters

The Verified Underreporter designation is awarded on any dimension where the following conditions are simultaneously met:

  1. The raw ratio exceeds 1.0 on its own evidence. Observed performance is better than claimed, not merely equal.
  2. The pattern is multi-instance. The gap is documented across at least three independent measurements or product lines — not a single favorable test.
  3. The understatement is structural, not accidental. There is evidence the vendor knowingly set conservative specifications — internal communications, design posture, or a consistent historical record — rather than simply underestimating their own system.
  4. No other dimension is below 0.9. A super-unity score on one axis cannot be used to redeem a deficiency on another. This rule preserves the core principle of the geometric mean: a chain is only as strong as its weakest link.

When all four conditions hold, the score is still computed with the clip — the aggregate remains in [0,1] — and the finding is published with the Verified Underreporter designation, the qualifying dimensions, and their raw pre-clip ratios (e.g. “E raw 1.12, published 1.00”). The reader gets the full information; the index keeps one definition.

What the designation means — and does not mean

The designation does not mean the system is perfect. It means the vendor's claims are structurally conservative relative to what they actually deliver. The organization tells you less than the truth in their favor. In a landscape where most published specifications are optimistic, aspirational, or outright wrong, a vendor that systematically underreports is a meaningful outlier. The designation records that outlier status explicitly rather than relegating it to a footnote.

It is rare by design. Verified Underreporter status is not awarded for beating a spec once, or for modest outperformance within normal manufacturing variance. It requires the kind of documented, intentional conservatism that an organization chooses as a brand posture and executes consistently. Most vendors will never qualify. The ones that do earn a designation that cannot be confused with “barely passing.”

Why not publish the super-unity number itself, as earlier versions permitted? Because MFid’s claim to be an index — and SDI’s claim to be the Software Defined Index — depends on one bounded scale every reader interprets identically. A published 1.12 invites comparison against a 0.94 as if the two lived on the same axis; they do not — one measures fidelity, the other generosity. The raw ratios are in the annotation for any reader who wants them; the scale stays [0,1].

Evidence tiers — one name, one definition

Every published MFid number carries an explicit evidence tier and a coverage percentage. The tiers, canonically:

  • Tier 1 — Measured Reality. What we observe in the client environment, under client load, on the client’s worst day. The verdict.
  • Tier 1E — Engineering Estimation. Where direct measurement of a subsystem is incomplete, the uncovered portion is scored by named inference method and labeled 1E. The inference method must be explicitly stated in the published finding. A Tier-1E sub-score is not a Tier-1 score; it is reported separately in the coverage label.
  • Tier 2 — Published Specification. What the vendor put in writing. The claim being tested. Example: NVMe latency vs. datasheet; ISP bandwidth vs. contract; cloud uptime vs. SLA.
  • Tier 3 — Scientific Calculation. What physics, mathematics, or architectural law allow. The ceiling no claim can exceed. Example: thermal throttling derived from TDP; channel capacity bounded by Shannon’s theorem.

v2.1 renumbering (2026-05-28). This is a renumbering, not a methodology change. The three categories and their definitions are unchanged. No published scores moved as a result. Tier 1 is now Measured Reality because that is where reader intuition puts the strongest evidence — the prior numbering (T1 = Scientific Calculation) inverted that intuition. The sub-label for incomplete measurement moves with the tier and is now 1E.

The naming history. Earlier versions of the site used “Measured Reality” on the homepage and “Engineering Estimation” on the manifesto for the same tier. That was a finding against ourselves. The reconciliation is above: one tier (Tier 1 — Measured Reality) with an explicit sub-label (1E — Engineering Estimation) only when the measurement is incomplete.

The coverage rule. Every published MFid number carries the form:

MFid 0.81 (Tier-1 over 72% of subsystems; Tier-2 published spec for the remaining 28%)

A score with no tier label and no coverage percentage is not a published MFid. It is a draft.

Limitations we will not hide

  • The four dimensions are not commensurable in a strict measurement-theoretic sense. The geometric-mean aggregation is a policy choice — it penalizes weakness — and not a derivation. We have argued the choice above; we have not proved it is unique.
  • Integrity is a meta-level measurement. It depends on a specification existing, being current, and being readable. The rubric handles this by scoring spec clarity (Ispec) as a sub-factor, but the dependence is real.
  • MFid does not score what has no spec. Vendor lock-in, switching cost, strategic dependency, and category-creation risk are real business exposures with no published claim to test against. MFid is silent on them by design. A high MFid on a vendor you cannot leave is still a problem; that problem belongs in procurement and architecture review, not in this instrument.
  • MFid does not detect compromise that stays within spec. A breach that does not violate any published SLI — a low-and-slow data exfiltration under the throughput ceiling, an authorized credential used by the wrong human — will register as I = 1.0 because behavior matches specification. Security incident detection is downstream of MFid, not a substitute for it.
  • MFid does not measure novelty. A system that does something genuinely new — a category that did not exist when its spec was written — will score against an obsolete spec. The instrument will say “faithful to claim” while the claim has been overtaken by what the system is actually being used for. Spec drift is real and Integrity’s Idrift sub-score catches some of it; the deeper problem of category creation is not solved here.
  • The aggregate MFid is computed by SDCorp. The framework, the rubrics, the rollups, the publication cadence are ours. The audit-of-the-audit problem is unresolved by this page alone. Standards engagement, academic co-authorship, and third-party attestation of a sample engagement are all on the roadmap; none is currently engaged or scheduled. We will not list them as “in progress” until they are. Status — honestly — published on the live status page.

Constraint on public-facing targets

An MFid target’s claim must be publicly verifiable to qualify for a public-facing teardown or walkthrough. Privately contracted SLAs, NDA-bound specifications, and access-restricted documentation are out of scope for SDCorp’s public artifacts because the audit source is not falsifiable by the reader.

Private targets remain valid for client-internal engagements; the constraint applies only to the public-facing artifact (blog post, evergreen teardown page, walkthrough). This constraint exists because MFid’s value to the reader depends on the reader being able to verify the underlying claim independently.

Versioning and the Revision Ledger

This methodology is versioned. The current version is MFid Methodology v2.6.1 (2026-08-05).

The full revision history lives on its own site, the SDCorp Revision Ledger — every formula change, rename, and correction, recorded there before it takes effect here.

Open the Revision Ledger →

Want the formula applied to your stack?

Bring the spec, we bring the math. The number does the rest.

Request an Investigation
Celestial Mode: Night