On 11 September 2026 the conformance headline for our HTML to PDF engine fell from 94% to 93.10%. The headline fell because the denominator grew and got harder, not because rendering got worse. Three things changed at once: a scripted file now counts as its subtests instead of as one unit, anchor positioning came back into the denominator, and a tier judging computed values started reporting. The Rust PDF engine improved the same day — reftests went from 93.62% to 95.49% across 39 accepted runs. Both scales are below, on one dataset.
Why we changed the count
A subtest is one test() or async_test() reported by the Web Platform Tests harness, and one subtest can hold many assertions. wpt.fyi is the public dashboard that records browser runs of those tests. It counts a reftest — a test that renders a page, renders a reference and compares them — as one unit, and a scripted file as its subtests. Until 11 September we counted reftests only.
css/css-flexbox/abspos/flex-abspos-staticpos-align-self-vertWM-001.html states 56 expected values and 29 of them match. Counted as one file it is a single failure, and those 29 matches are recorded nowhere. Our corpus holds 26,325 reftest units and 57,046 units in total.
Every published conformance number we know of — wpt.fyi, Servo (the browser engine written in Rust), Blitz — counts the harness unit. A number on another scale cannot be laid beside any of them.
A print engine has a second reason. We have no JavaScript runtime, so 2,362 files state their expectations only from script and we cannot run them. On the old scale those files were absent from both halves of the fraction. They are now counted, named and shown with the reason: the third headline on the conformance page treats every unjudged file as a failure and stands at 86.74%.
What the new scale shows
A line-by-line comparison with wpt.fyi. The conformance page carries a coverage column: of the 34,511 tests wpt.fyi records under css/, our corpus judges 26,596, which is 77.1%; of its 218,751 subtests we judge 55,154, which is 25.2%. Snapshot: Chrome 155.0.8052.0, 11 September 2026. The two totals differ by 1,892 units: our corpus reaches outside css/ — html/rendering, html/semantics, mathml, svg and compat — and the snapshot covers css/ only.
Two additions to the denominator. Anchor positioning had been excluded as unshipped and came back: 185 of 249 tests and 518 of 614 subtests in this dataset. A tier — one family of tests with a unit of its own — began judging computed values, meaning getComputedStyle serialization as specified in CSSOM §6.7.2: 342 files, 3,266 of 3,846 subtests passing, and 460 subtests set aside because the serializer for those properties is not written yet. Both areas score below the corpus average.
Where the figures come from. Every number on the conformance page is read from the dataset built after the last accepted run — this page reads 20261001-0535-e08c90ab-d2601fe3 — and each test links to its own evidence: our render, the reference, and the pixels that differ.
The numbers as they stand
| Scale | Unit | Result | Dataset |
|---|---|---|---|
| Reftests only, as published in August | one reftest = one unit | 94% | 27 August 2026 |
| Reftests only, recomputed on the current engine | one reftest = one unit | 95.49% | 1 October 2026 |
| wpt.fyi scale — the headline now | reftest = 1, scripted file = its subtests | 93.10% | 1 October 2026 |
| By files | one file = one unit | 93.95% | 1 October 2026 |
| Score, as Servo defines it | per test: the share of its subtests that pass, where a reftest scores 0 or 1; the mean over all tests | 94.71 | 1 October 2026 |
| Everything unjudged as a failure | one file = one unit | 86.74% | 1 October 2026 |
Where the units come from — the four tiers
| Tier | What it judges | Units | Passing |
|---|---|---|---|
| Reftests | rendering, against the test's own reference | 26,325 | 95.49% |
| Layout assertions | element positions, one unit per element checked by check-layout-th.js | 17,094 | 95.09% |
| Parsing acceptance | which declarations survive the parser, one unit per parsing-testcommon.js assertion | 9,781 | 86.38% |
| Computed values | getComputedStyle serialization, CSSOM §6.7.2 | 3,846 | 84.92% |
Reftests — the part the old number measured — stand at 95.49% on this dataset, above the 94% published in August. The accepted runs behind that move are listed in the run history on the conformance page, each with the tests it turned green.
Print modules are never excused. A failing test leaves the denominator only when fewer than three of the four browser engines pass it and its specification is below Candidate Recommendation, the W3C stage at which a specification is considered stable. That rule never applies to print modules: css-break, css-page and css-multicol decide where a page ends, which is the part of CSS a PDF engine exists to get right. Controlling those page breaks from a document is covered in the documentation.
What moves the number next. 2,362 files need a JavaScript runtime to state their expectations, and they are the largest block we cannot judge. A testharness.js runner now exists: on a sample of 50 such files it produced verdicts for 24, and the subtest names and counts matched Chrome on all 24; the rest wait on CSSOM. Its results reach the page once they pass the acceptance gate. The per-directory table has a row each for CSS Grid and Flexbox.
Frequently asked
Why did the pass rate go down?
Three changes to the denominator landed on 11 September 2026: a scripted file counts as its subtests rather than as one unit, anchor positioning returned after having been excluded, and a tier judging computed values started reporting. Rendering improved the same day: reftests stand at 95.49% against the 94% of the old headline, across 39 accepted runs.
How is a subtest counted?
A subtest is one test() or async_test() reported by the harness. The layout tier counts one unit per element checked by check-layout-th.js, and the parsing tier one unit per assertion in parsing-testcommon.js, because those are the units wpt.fyi records for those files.
Do you drop tests you fail?
A test we pass is always judged. A test we fail leaves the denominator only under three conditions at once: fewer than three of four browser engines pass it, its specification is below Candidate Recommendation, and it is not a print module. The rule and the counts are in the methodology.
Can I check a single test myself?
Every test has a page with our render, the reference and the pixels that differ, and its source opens in the playground.