Mizan

Method & Evidence

Two things we measured about our own instruments came back badly, and this page reports them. Where we failed a threshold we set ourselves, we say so, and we say what we stopped claiming as a result.

What we measured, and what it showed

The framework is only worth anything if two people applying it to the same material land in the same place. The pilot tested that: 27 decision choices from two courses, difficult conversations and conflict management, with the authored judgment stripped out, coded for cylinder and mode by two raters working independently and blind to each other.

Two independent raters agreed on the joint code for 20 of 27 choices (κ = 0.68, 95% CI 0.48–0.86) against coding manual v0.1.

We take the verdict from the lower bound of the bootstrap interval and treat 0.61 as the threshold. The lower bound was 0.477. The pilot did not clear the bar we set for it. The looser reading belongs here too: the point estimate of 0.68 lands inside the band our own spec table calls usable internally, revise and re-run. Both readings prescribe the same next action, so we report the stricter one. The item-count bar we stated before the run was itself computed on the wrong chance-agreement figure and had to be corrected afterwards.

The two raters were Claude and Gemini — two language models, not two people. Our own protocol says the figure requires two competent humans coding independently, and that a model standing in for a human rater is not inter-rater reliability. So this is a rehearsal of the procedure rather than the result it is meant to produce. We publish it with that limit attached rather than wait for a better number.

Three further findings constrain what we may say:

  • Remove every Cylinder 7 item and the 21 choices left agree at κ = 0.87 (CI 0.68–1.00). That number carries a confound this pool cannot break: every Cylinder 7 item sits in conflict management, the course one rater did not author, so dropping Cylinder 7 and dropping that course are nearly the same operation. We do not publish "Cylinder 7 is our weak cylinder." It may be true; this data cannot show it. Cylinder 7 stays out of reports either way.
  • Four of the seven cylinders — Safety, Belonging, Growth and Meaning — appear nowhere in the 27 choices, because both courses are interpersonal-conflict content. This catalogue supports a two-cylinder read today, not a seven.
  • Each rater used the code excess exactly twice, against a 20 percent floor our spec sets, and never once on the same item. Agreement on excess is zero. That is the code that would tell you which of the two directions a cylinder is failing in, and it is why nothing we publish reports that split.

The manual is now at v0.3; the reliability figure quoted above is from v0.1, and should never appear without it. Every revision traces to a specific disagreement in the pilot. The two raters then wrote divergent v0.2 revisions from the same disagreement list and ruled opposite ways on three items, so neither v0.2 was canonical; the framework's owner adjudicated those three on 9 August 2026, and v0.3 supersedes both. No κ has been computed against v0.3, and the 27 pilot items are spent for that purpose — a second-round figure has to come from material neither rater has seen.

One re-run exists and it is not a reliability figure. On 9 August 2026 the second rater re-coded the same 27 items against its own v0.2 and returned joint κ = 0.81 (CI 0.63–0.95), clearing the bar. It does not count, for two reasons we would rather state than have found: that rater had computed the v0.1 result itself, which required reading the first rater's full sheet, so the second pass was not blind; and it sets a v0.2 pass against a v0.1 pass. The most it may be called is manual-guided convergence of 0.81, non-blind. The number we quote as reliability is still 0.68, from v0.1.

The scores we retired

The free structure scan used to return a health score, a drag score and a strategy-alignment score. Drag was the variance of manager spans around the mean, times ten, plus five for each named bottleneck, capped at 100. Health was 100 minus drag.

Variance dominated the bottleneck term, so the metric measured how uniform spans were and called that health. Uniform pathology scored well. Measured over the live endpoint: 60 people all reporting to one manager scored 95. Five managers with 40 reports each scored 75. An ordinary organization with spans of 1 through 6 and no overloaded manager scored 67. The worst structure constructible scored 28 points healthier than an unremarkable one.

There is no better formula to put in its place. Organizational drag is a phrase from a Bain book, defined there and given no published calculation anywhere. A single composite index fails the OECD/JRC handbook's requirements: weights in such an index are value judgments, a compensatory aggregate can disguise a serious failing in one dimension, and uncertainty analysis is step 7 of 10 rather than an optional extra. And no source that discloses its data claims an optimal span of control — the one study that tested the hypothesis directly rejected it, Meier and Bohte finding across 678 Texas school districts that schools with average spans of control are less productive. Span varies roughly twentyfold between sectors: 3 to 5 across the Australian Public Service, against a median of 46 for inpatient nurse managers.

What the scan returns now is a description with the arithmetic printed beside it — span distribution rather than an average, depth profile, management ratio, single-report managers, structural anomalies, the four Krackhardt outtree axioms, and a section on what a reporting chart cannot tell anyone. See the structure analysis.

Your answers do not go to your employer

An individual's read belongs to that individual. This is a different promise from the retention notice on the survey page: retention says when a record disappears, this says who can read it while it exists.

The promise exists because of what the instrument is. A read of how a workplace actually feels is worth running only if people answer accurately, and people answer accurately only when the answer cannot be used against them. An assessment that can be traced back to the person who filled it in measures caution rather than culture. The firewall is the condition under which the survey returns anything true, not a courtesy attached at the end.

The public survey takes no name, no email, no account and no employer login. One row is written; it holds your answers and whatever company details you typed as context, and it holds no name, email, account id or employer id, because the table has no column for any of them. The browser that submitted it receives a private link, and that link is the only way to open the report. The row is deleted automatically within 24 hours by a purge job that runs hourly and again at startup.

One thing outlives that row, and this page would be dishonest if it did not say so. When your report is built, the fourteen scores it shows — the withheld and overextended figures for each of the seven domains, plus the values-match and engagement answers — are written to a second table and kept. That row carries no name, no email, no company details, not the link to your report and not your individual answers; it does not record a time of day, only a date, so it cannot be matched against a server log. It exists because the balance threshold has to be set from a real distribution rather than chosen, and until it is, every report on this site says the classification is pending. The first design kept nothing at all, which sounded stronger and meant the threshold could never be calibrated: five days after the survey opened, three people had taken it and all three response sets had already been deleted.

In the signed-in organizational product the same rule runs from the other side. An individual report can be opened by the person it is about and by nobody else — role is not a way in, and the access check takes no role argument at all, so a future caller cannot pass one. Organizations receive aggregates, and an aggregate is withheld entirely below five responses rather than shown with a caveat. That floor is about subtraction rather than noise: four responses today and five tomorrow would make the fifth person's answer the difference between two published averages.

We wrote this section before the code was true, found three places where it was not, and fixed them first. Until 11 August 2026 any administrator in a tenant could open a named employee's individual culture report, and the organization-level endpoint returned raw individual rows.

The rule the recommendations follow

The two-sided shape of the recommendations is older than the framework. In the twenty-second book of the Ihya’ Ulum al-Din, al-Ghazali treats character the way a physician treats a body: a trait held too little is corrected by encouraging its opposite — but only so far, because the correction itself overshoots into the opposite disorder, the way treating cold with heat, pressed too far, ends in fever. Avarice is cured by giving, but not to the point of squandering, which is a disorder of its own. A recommendation that pushes in one direction with no point at which more stops helping is exactly the error that rule warns against, and it is why a report that says “bring more of this value” without saying where more turns to harm is one we treat as unfinished.

The picture the idea travels on is an ant dropped into a ring heated at its rim: it runs from the heat on every side until it settles at the center, not as a compromise between two goods but as the one point not on fire. That is what a balance is for — the mean is a place a value can miss in two directions, not a maximum to drive toward. It is the same commitment that keeps this page from reporting a direction of failure it cannot yet measure.

Paraphrased, not quoted, from al-Ghazali, Ihya’ Ulum al-Din XXII (On Disciplining the Soul), tr. T. J. Winter, Islamic Texts Society; the doctrine of the mean it applies is Miskawayh’s.

What we do not claim

  • No Cylinder 7 finding in any report. The code is not yet reliable, and this pool cannot establish that Cylinder 7 is the reason.
  • No split between the two failure directions in any report. Two raters agreed on zero excess codes. The framework names both directions; no instrument here tells you which one you are in.
  • No seven-cylinder culture read from the course catalogue. Four of the seven do not appear in it.
  • We do not call the pilot publishable, and we do not quote κ ≥ 0.80 as its result. The lower bound was 0.477, and the non-blind 0.81 re-code is not offered in its place.
  • Nothing is claimed for manual v0.3. It has no κ of its own yet.
  • No structure health score, drag score or strategy-alignment score.
  • No target or ideal span, and no target number of layers.
  • Nobody on a reporting chart is called a bottleneck. A bottleneck is a property of flow, and a reporting tree contains no flow. A high span is a high span; a long sole-reporting chain is that.
  • No claim that fewer layers means faster decisions. We looked for a study supporting it and found none.
  • No claim about information flow, silos, collaboration or influence read from reporting lines.
  • No benchmark we cannot reproduce from a disclosed sample.
  • No report produced when an analysis fails. If an assessment cannot complete, it tells you it failed and returns nothing — there is no template, no estimated score and no partial reading standing in for a result.

How to check us

The pilot is deterministic: its command, bootstrap count and seed are recorded with the result, so the same sheets give the same interval every time. The rater sheets, the tool and the machine-readable output are files; ask and we will send them. The second rater's re-code sheet is the one exception — it was never returned, and what we hold is a reconstruction from that rater's own report, marked as such. If you see 0.68 quoted anywhere without v0.1 beside it, that is our error.

The sources behind the retirement are public and dated, and each was fetched rather than recalled: Meier and Bohte, the OECD/JRC handbook on composite indicators, Bain's own page, and Gallup's January 2026 data on 16,442 US managers, where mean team size rose from 8.2 to 12.1 while the median held at 5 to 6.

The retirement is enforced in code rather than in policy. The endpoint that once returned the scores carries the old formulas and the measured degeneracy in a comment, so anyone proposing a headline number again meets the reason it was removed. If you find a figure on this site that does not trace back to something here, tell us and we will source it or take it down.