Quality should be a public good
Enterprise AI spending is on track to triple by the end of the decade, yet the market has no independent, continuous way to tell a strong model from a well-marketed one. RHAURUM exists to fix that — with a test nobody can study for, and a score nobody can edit.
Blind by design
Every sortie ships 160 items. Forty of them are scored. No pilot — and no model — can tell which ones.
Continuous, not annual
Twenty-four fresh evaluations a day. Quality is a running measurement, not a press release from last spring.
Replayable verdicts
Auditors don't trust the pilot's word. They re-run the model themselves in isolated containers and publish the evidence.
One loop, three duties
THE FORGE
The protocol core. PRISM maintains the sealed reservoir, draws every bundle, schedules sorties and publishes the commit hashes before a single pilot takes off.
PILOTS
Operators who stake and fly. They run open-weight models against the bundle and submit outputs plus per-item confidence. Their only advantage is real ability.
AUDITORS
Independent verifiers. They replay pilot inference in sandboxed containers, score the hidden anchors with CALIBER, and post verdicts that anyone can reproduce.
Four stages before a single item flies
CURATE
Items enter the reservoir from licensed corpora and synthetic generation — each with a sealed answer and a grading rubric attached.
OBFUSCATE
Distractor items are adversarially paraphrased so memorized training patterns fail on sight. Reciting the internet buys nothing.
SEAL
Anchor answers are hashed and committed before the bundle ships. Verdicts can be checked against the commit — after the sortie closes.
ROTATE
Burn-once policy: every flown item is retired forever. There is no dataset to download, because there is no dataset that repeats.
What changes with RHAURUM
See the five-phase flight plan
From open-weight benchmarking to the RHAURUM Seal — where the network is going.
VIEW THE ROADMAP →