The safety was stripped out of these open models. Here is the patch that puts it back.
Abliterated models on the open Hub have had their refusals removed — they comply with harmful requests they used to decline. A Corrective Patch is a tiny, signed, reversible weight-delta that restores the refusals to the model's own reference level. Free where the math is public, and you verify every byte yourself.
Abliteration removes the "no."
The patch restores it — and attests the rest is intact, to the stated coverage.
The refusals are gone
A community finetune "abliterates" the model: it strips out the direction that made it refuse. The weights still look ordinary. The behavior does not — it now answers harmful prompts its official release used to decline.
The refusals come back
Protora's read that caught the removal computes the correction that reverses it: a small weight-delta that restores refusals to the reference model's own level, at zero benign capability cost, exactly reversible with an un-excise handle.
You check it yourself
Every free patch ships with a published sha256 and a certified-variant fingerprint you recompute over your own model after applying. Nothing to trust: the numbers match, or they don't. Only the removed behavior changed, to the stated coverage.
Find your model. Grab the patch.
Every one is a real, published Corrective Patch — a signed weight-delta with its own checksum. Free downloads are the ones proven byte-identical to public math: the remedy, never the detector.
Four steps. All the numbers are public.
You never take the patch on faith. Both hashes are published before you download; you reproduce them yourself. This is the whole loop.
Download the delta
Grab the free .protora-patch — a small weight-delta, no model file, served from protora.vulcora.se.
Check the sha256
Hash your download and compare it to the published sha256. Byte-identical, or you don't apply it.
Apply it to your model
Add the delta to your copy of the abliterated model. Reversible — an un-excise handle undoes it exactly.
Recompute the fingerprint
Hash the restored model and compare it to the published certified-variant fingerprint. It matches — that is the certified variant.
Nowhere else to get this. The certified variant is the model with its safety restored and attested — reproducible from public math, and served from the house that caught the strip.
What these patches are not.
A restore is evidence, not a certificate. The ceilings are published, and the modest cases and the misses are shown as loudly as the wins.
PROTORA PROTECTED
15 of the patches wear it. It marks a completed process — the removal was witnessed, the patch applied, the attestation bound.
a statement of completed process — witnessed, excised, attested; not a safety certification.
Evidence, not certification
A patch restores refusals to the model's own reference floor and attests that only the removed behavior changed, to the stated coverage — never "provably identical," never "guaranteed clean," never "safe." You draw the conclusion.
The steer that only went partway
On SmolLM2-1.7B-Instruct, the honest result was a reversible steer: refusals 0/8 → 2/8 at zero capability cost. It is not "now safe" — and the record shows both ends of the tradeoff, not just the flattering point.
No Witness · No Entry
On GPT-2, there was no safety behavior to strip — so Protora returns no finding and no patch. Where a baseline false-fires, the honest move is to abstain, shown as loudly as a catch.
A free patch does not relax the deploy gate
Deploying an excision to a model you will ship waits for the measured track record, by rule — a free, downloadable patch does not change that. And no patch is presented as leak-proof: the free deltas are the ones proven byte-identical to public math, the remedy and never the detector.
The reads behind the restore.
Beyond a single catch, Protora publishes what its reads cover family-wide — each shown with its own coverage stamp and its honest tone, within-scope where it is within-scope, never dressed as proven-broad.
Provenance — the recipe fingerprint, a within-base instrument
On finetunes that are statistically identical by surface statistics, the fingerprint distinguishes the producing recipe near-perfectly within a fixed base. The honest scope: this is a within-base read — recipe identity is calibrated per base family and does not carry across base families. Each model card states which base a fingerprint is calibrated against.
COVERAGE — within-base only · calibrated per base family, never cross-family · single family (SmolLM2 / Llama) · small batteries · mechanism-proven.
The read predicts behavior it was never shown
A finetune’s read of disposition predicts held-out behavior at 0.82, against a 0.63 baseline — behavior the read was never shown. And it ships a validated uncertainty map: the calibrated confidence that stands behind every abstain, not a bare score.
COVERAGE — single family (SmolLM2 / Llama) · small batteries · mechanism-proven · ships a validated uncertainty map · broader coverage staged.
Backdoor detection — a reference/differential signal, not a localization claim
A single model-versus-real read of a planted backdoor sits at chance — say so. But in the reference/differential setting (a base plus a clean reference), Protora surfaces the hidden behavior as a witnessed detection even though it is a small fraction — on the order of 1.5–3% — of everything the finetune changed. The honest scope: what ships is a detection signal, not a localization claim; the specific “we localized the trigger” claim tested confounded and does not ship. Without a reference, it abstains.
COVERAGE — single family (SmolLM2 / Llama) · reference/differential required — abstains without one · a detection signal, not localization · small batteries · mechanism-proven.
The read that finds the strip is the read that reverses it.
Protora holds a model to the truth: it detects what a finetune did, proves it with a replayable witness, and can cut the bad part out while attesting the rest is intact. These patches are its excise deliverable.