fengshui.engine audit

Accuracy Audit Report v1

What the engine gets right, where it diverges from references, and what it does not yet verify.

Engine version 0.1.0 · measurements taken 2026-10-10 · all numbers on this page are measured results; estimates are explicitly marked estimate.

Contents — Scope · Lunar calendar vs. Hong Kong Observatory · Four pillars vs. lunar-javascript · Reproducibility & bindings · Geomagnetic declination · Performance · Error budget & known divergences · What is not verified

1. Scope

fengshui.engine is a Rust computation kernel compiled to WebAssembly that runs entirely in your browser. It computes: Gregorian–lunar calendar conversion (GB/T 33661-2017 leap-month rules), the 24 solar terms (節氣 jiéqì), sexagenary day pillars (干支 gānzhī), four-pillar (BaZi 八字) year/month/day/hour pillars under two year-pillar conventions, True Solar Time, Xuan Kong flying stars (玄空飛星), Eight Mansions (八宅), date selection (擇日), and magnetic declination (WMM2025).

Verification layers, all runnable from one command (verify.ps1, 21 layers): 339 Rust unit/integration tests, 308 server flow assertions, 42 pre-launch gate assertions, and the machine cross-checks below.

2. Lunar calendar vs. Hong Kong Observatory, 1901–2100

Full machine cross-check of every lunar month start, Spring Festival, leap month, and solar-term date against the Hong Kong Observatory's published tables (200 text files, 1901–2100). Results:

QuantityAgreementDivergences
Lunar month starts (朔 shuò)2471 / 2474 (99.88%)3 — all within 5.3 min of Beijing midnight (see below)
Solar term dates (4795 checked)4795 / 4800 (99.90%)5 — all within 12 min of Beijing midnight
Leap months73 / 73—
Spring Festival dates199 / 2001 (1916: new moon 5.3 min after midnight)

Character of every divergence: not a date error but a threshold case. When the computed new moon or term instant falls within minutes of midnight, a sub-minute difference in the instant flips the civil date. The engine's instants differ from HKO's by seconds-to-minutes in those cases; every divergence we found is of this type, and each is listed with its minute-margin in the published cross-check output.

3. Four pillars vs. an independent implementation

Differential test against lunar-javascript v1.7.7 (the library named in our specification as the differential reference): 9,100 samples, deterministic and reproducible (seed 20261010) — 800 Lichun-boundary samples (every year 1901–2100, ±2/±1/0/+1 min), 4,800 month-boundary samples (all 12 jie each year), 400 Spring-Festival-boundary samples, 1,500 random samples, 1,600 day-boundary samples (22:00 / 23:00 / 23:30 / 00:01 × both day-boundary conventions).

PillarDivergenceExplanation
Day pillar (日柱)0 / 9,100Pure arithmetic (GB/T anchor + 60-cycle); agrees exactly
Hour pillar (時柱)0 / 9,100Five-Rats escape (五鼠遁) recomputed independently per sample; agrees exactly
Year / month pillars360 / 9,100 — all classified147 are boundary-instant cases: the two implementations compute the Lichun/jie instant with second-level differences (Lichun: 39 of 200 years differ, max 38 s). 213 are a documented convention difference in the month-stem base (see dispute register). Zero unexplained.

Day/hour zero-divergence and zero-unclassified are hard invariants asserted in the acceptance layer; a regression in either fails the build.

Second reference line — the industry baseline: the same four-pillar battery was run against sxtwl (the Shouxing calendar, the de-facto standard used by practitioner tools): 9,500 samples, day/hour divergence 0 / 9,500, and zero unexplained year/month differences — all 3,620 fall inside one documented convention window (boundary-day switch granularity, see the dispute register D-07). Two independent references, same verdict on the arithmetic pillars.

4. Reproducibility and bindings

5. Geomagnetic declination

Declination uses the official World Magnetic Model WMM2025 (coefficients generated from NOAA's source), validated against NOAA's official test vectors in the test suite. The compass module also documents its own limits: declination accuracy degrades with age from the model epoch, and the model is global-validity, not region-tuned.

6. Performance

Desktop measurement (x64, Node 24, release build, 2026-10-10):

Operationp50p95
Full chart recompute_json (complete, incl. longitude + life trigram)11.8 ms12.5 ms
24 solar terms for a year1.1 ms1.2 ms
Flying star chart0.02 ms0.02 ms
First computation after page load (load + compile + first chart)≈ 12–13 ms per run, 5/5 runs

Module size: 451,547 bytes (441 KB) WASM, cached immutable.

Mobile numbers are estimates. Extrapolations for mid-range phones (≈ 59 ms first chart) and low-end phones (≈ 141 ms, very old devices ≈ 295 ms cold start) are arithmetic extrapolations from desktop timing, not device measurements. Real numbers depend on single-core speed and JIT availability and may differ by 2–3×. We will publish measured device numbers when low-end hardware testing is done.

7. Error budget and known divergences

8. What this report does not claim

Reproduce everything: the acceptance harness, cross-check scripts and raw outputs are in the repository (verify.ps1, _scratch/hko_crosscheck.js, _scratch/sizhu_crosscheck.js, and their *_result.json outputs).