stackrank.ing

Calibration

A rating on this site describes a stack's position relative to the other stacks reviewed in the same quarter. It does not describe the stack on its own. This page explains how that works and why it is done this way.

What a rating is

Each stack is scored on five dimensions from 1.0 to 5.0. Those scores are multiplied by the weights in force for the quarter and summed, which gives a composite. The composite is an absolute measurement and it is published on every stack page.

The rating is a separate thing. Once every composite in the quarter is known, the cohort is sorted from highest to lowest and the target distribution is applied to the sorted list. The stack in first place receives whatever rating the top bucket holds. The stack in last place receives whatever the bottom bucket holds. Nothing about that depends on how high or low the composites happen to be that quarter.

This is the part people find surprising, so it is worth stating plainly. If every stack in the cohort improved substantially, the ratings would not change at all. The same number of stacks would land in each bucket, in the same order, because the distribution is a property of the cohort rather than of the work.

Why the distribution is fixed

The alternative is a fixed threshold: decide in advance that a composite above 4.0 is a top rating, and award it to everybody who clears the bar. That was the original design and it was abandoned after two quarters, for two reasons.

The first is that the thresholds drift. Reviewers who score twelve stacks in a sitting score the later ones more generously than the earlier ones. The absolute numbers move without the pancakes moving. A distribution applied after the fact is immune to that, because it only uses the order, and the order is much more stable than the values.

The second is that a threshold system produces no information in the years when everything is good. A quarter in which every stack clears 4.0 tells the reader nothing about which stack to make on a Sunday. The forced distribution guarantees that the published record separates the cohort every quarter, which is the only reason a reader would consult it.

How bucket sizes are calculated

The targets are proportions and the cohort is a whole number, so the two rarely divide cleanly. Bucket sizes are resolved by the largest remainder method: multiply each target by the cohort size, take the whole part, then hand out the seats left over to the buckets with the largest fractional parts, in the order the buckets are declared.

The 2026-Q3 cohort holds 11 stacks. The arithmetic is below and it is the arithmetic the site actually runs. The build fails if the allocated column does not sum to the cohort size.

Allocation for 2026-Q3, cohort of 11.

BucketTargetExactWhole partAllocated
Greatly Exceeds5%0.5500
Exceeds15%1.6512
Meets60%6.6067
Meets Most15%1.6511
Does Not Meet5%0.5501
Total100%11.00811

The two rules

Two constraints are applied after the allocation, and they are the only points at which the boundaries are moved by hand rather than by arithmetic.

The bottom bucket must contain at least 1 stack every quarter. If the rounding leaves it empty, the boundary moves up until it is occupied. The reasoning is that a review programme which never records a poor result is not reviewing anything. Somebody has to be last, and the record is more useful if it says who.

The top bucket holds at most 1 stack. There is no corresponding minimum, which is why the top bucket is sometimes empty. In 2026-Q3 the target of 5 per cent resolved to 0.55 stacks and the seats left over went to buckets with larger fractional parts, so Greatly Exceeds was not allocated. The highest composite in the cohort this quarter is Japanese Soufflé Stack at 4.5, rated Exceeds.

That outcome is not an error and it is not adjusted. A bucket that is awarded in every quarter regardless of the cohort is not a top bucket, it is a formality, and the programme has no way to say anything with a formality.

What follows from this

A composite can fall while every dimension score holds or improves. This happens whenever the weights are revised, because the composite is a weighted sum and the weights are part of it. Two stacks in the current cohort are in exactly that position, and their reviews say so in the review body rather than in a footnote.

A rating can fall while the composite holds. This happens whenever a stronger stack enters the cohort, because the buckets are filled by rank and a new stack above you moves you down one place in the sorted list. Nothing about your stack changed. The list got longer above you.

Both of these are the system working. Neither is a signal about the stack, and the calibration note on each stack page exists specifically to say which of the two is responsible for the movement that quarter, so that the reader does not have to infer it.

Objections

The bottom rating is unfair to a stack in a strong cohort. It is, and this is accepted rather than denied. A stack rated Does Not Meet in a strong quarter would have been rated higher in a weaker one. The response is that the composite is published alongside the rating on every page, and the composite is not relative. A reader who wants the absolute measurement has it.

The weights are revised too often. Three quarters, two revisions. Each revision is recorded with the quarter it takes effect from and is never applied retroactively, so an archived quarter can always be read with the weights that produced it. The revision history is on the rubric page and every archived quarter carries the weights it was calibrated under.

Ratings should reflect the work, not the cohort. The work is the composite. The rating is the cohort. The site publishes both, on the same row of the same table, and does not average them into a single number, because averaging them would destroy the only thing that makes either of them readable.

What calibration does not do

Calibration does not determine level. Level describes expected scope and is assessed separately, on the leveling page. A stack can hold the lowest rating in the cohort and keep its level, and Boxed Mix, Just Add Water has done so for three consecutive quarters.

Calibration does not trigger removal from the cohort. Removal is a separate decision made against the coverage of the review programme, and the one stack removed to date left for a reason unrelated to its rating. The record is on the alumni page.

Calibration does not carry forward. Each quarter is calibrated from that quarter's composites against that quarter's cohort. Prior ratings are not an input, which is why a stack can move two buckets in one quarter without anything having gone wrong.