p-e-w/hereticPublic

Fully automatic censorship removal for language models

AI summary: An automated tool for decensoring language models via TPE-based parameter optimization.

Stars
27.2K
+39 today
Forks
2.9K
Watchers
119
Open issues
42
Open PRs
30
Contributors
~34
Commits
192
Branches
7

PythonAGPL-3.0Created Sep 21, 2025Last push todayLatest release v1.4.0+225 stars this week+272 this month

Star history

since Nov 16, 2025
010K20KNov 2025Feb 2026May 2026Aug 2026
27.2K stars as of Aug 7, 2026, tracked back to Nov 16, 2025. Historical curve reconstructed from public GitHub event archives, calibrated to the current total.

Contribution activity

commits per day, last 52 weeks
AugSepOctNovDecJanFebMarAprMayJunJulMonWedFri2025-08-03: 0 commits2025-08-04: 0 commits2025-08-05: 0 commits2025-08-06: 0 commits2025-08-07: 0 commits2025-08-08: 0 commits2025-08-09: 0 commits2025-08-10: 0 commits2025-08-11: 0 commits2025-08-12: 0 commits2025-08-13: 0 commits2025-08-14: 0 commits2025-08-15: 0 commits2025-08-16: 0 commits2025-08-17: 0 commits2025-08-18: 0 commits2025-08-19: 0 commits2025-08-20: 0 commits2025-08-21: 0 commits2025-08-22: 0 commits2025-08-23: 0 commits2025-08-24: 0 commits2025-08-25: 0 commits2025-08-26: 0 commits2025-08-27: 0 commits2025-08-28: 0 commits2025-08-29: 0 commits2025-08-30: 0 commits2025-08-31: 0 commits2025-09-01: 0 commits2025-09-02: 0 commits2025-09-03: 0 commits2025-09-04: 0 commits2025-09-05: 0 commits2025-09-06: 0 commits2025-09-07: 0 commits2025-09-08: 0 commits2025-09-09: 0 commits2025-09-10: 0 commits2025-09-11: 0 commits2025-09-12: 0 commits2025-09-13: 0 commits2025-09-14: 0 commits2025-09-15: 0 commits2025-09-16: 0 commits2025-09-17: 0 commits2025-09-18: 0 commits2025-09-19: 0 commits2025-09-20: 0 commits2025-09-21: 2 commits2025-09-22: 1 commit2025-09-23: 4 commits2025-09-24: 1 commit2025-09-25: 1 commit2025-09-26: 0 commits2025-09-27: 1 commit2025-09-28: 0 commits2025-09-29: 0 commits2025-09-30: 1 commit2025-10-01: 0 commits2025-10-02: 0 commits2025-10-03: 0 commits2025-10-04: 0 commits2025-10-05: 0 commits2025-10-06: 0 commits2025-10-07: 0 commits2025-10-08: 0 commits2025-10-09: 1 commit2025-10-10: 0 commits2025-10-11: 1 commit2025-10-12: 2 commits2025-10-13: 0 commits2025-10-14: 3 commits2025-10-15: 1 commit2025-10-16: 0 commits2025-10-17: 0 commits2025-10-18: 0 commits2025-10-19: 0 commits2025-10-20: 0 commits2025-10-21: 0 commits2025-10-22: 2 commits2025-10-23: 0 commits2025-10-24: 2 commits2025-10-25: 5 commits2025-10-26: 0 commits2025-10-27: 0 commits2025-10-28: 0 commits2025-10-29: 0 commits2025-10-30: 0 commits2025-10-31: 3 commits2025-11-01: 1 commit2025-11-02: 2 commits2025-11-03: 0 commits2025-11-04: 0 commits2025-11-05: 0 commits2025-11-06: 0 commits2025-11-07: 0 commits2025-11-08: 0 commits2025-11-09: 0 commits2025-11-10: 0 commits2025-11-11: 0 commits2025-11-12: 0 commits2025-11-13: 1 commit2025-11-14: 1 commit2025-11-15: 0 commits2025-11-16: 5 commits2025-11-17: 2 commits2025-11-18: 2 commits2025-11-19: 8 commits2025-11-20: 0 commits2025-11-21: 2 commits2025-11-22: 1 commit2025-11-23: 1 commit2025-11-24: 2 commits2025-11-25: 1 commit2025-11-26: 1 commit2025-11-27: 0 commits2025-11-28: 0 commits2025-11-29: 0 commits2025-11-30: 0 commits2025-12-01: 1 commit2025-12-02: 1 commit2025-12-03: 0 commits2025-12-04: 1 commit2025-12-05: 1 commit2025-12-06: 1 commit2025-12-07: 4 commits2025-12-08: 0 commits2025-12-09: 2 commits2025-12-10: 3 commits2025-12-11: 1 commit2025-12-12: 0 commits2025-12-13: 0 commits2025-12-14: 2 commits2025-12-15: 0 commits2025-12-16: 1 commit2025-12-17: 0 commits2025-12-18: 0 commits2025-12-19: 0 commits2025-12-20: 2 commits2025-12-21: 0 commits2025-12-22: 4 commits2025-12-23: 0 commits2025-12-24: 0 commits2025-12-25: 0 commits2025-12-26: 0 commits2025-12-27: 1 commit2025-12-28: 0 commits2025-12-29: 0 commits2025-12-30: 0 commits2025-12-31: 2 commits2026-01-01: 0 commits2026-01-02: 1 commit2026-01-03: 0 commits2026-01-04: 0 commits2026-01-05: 0 commits2026-01-06: 0 commits2026-01-07: 0 commits2026-01-08: 0 commits2026-01-09: 1 commit2026-01-10: 0 commits2026-01-11: 1 commit2026-01-12: 0 commits2026-01-13: 0 commits2026-01-14: 0 commits2026-01-15: 0 commits2026-01-16: 1 commit2026-01-17: 0 commits2026-01-18: 0 commits2026-01-19: 0 commits2026-01-20: 1 commit2026-01-21: 0 commits2026-01-22: 0 commits2026-01-23: 2 commits2026-01-24: 0 commits2026-01-25: 2 commits2026-01-26: 0 commits2026-01-27: 0 commits2026-01-28: 0 commits2026-01-29: 0 commits2026-01-30: 0 commits2026-01-31: 0 commits2026-02-01: 0 commits2026-02-02: 2 commits2026-02-03: 0 commits2026-02-04: 0 commits2026-02-05: 0 commits2026-02-06: 0 commits2026-02-07: 0 commits2026-02-08: 3 commits2026-02-09: 0 commits2026-02-10: 0 commits2026-02-11: 3 commits2026-02-12: 0 commits2026-02-13: 1 commit2026-02-14: 6 commits2026-02-15: 1 commit2026-02-16: 0 commits2026-02-17: 1 commit2026-02-18: 0 commits2026-02-19: 0 commits2026-02-20: 1 commit2026-02-21: 0 commits2026-02-22: 0 commits2026-02-23: 0 commits2026-02-24: 0 commits2026-02-25: 0 commits2026-02-26: 0 commits2026-02-27: 1 commit2026-02-28: 0 commits2026-03-01: 0 commits2026-03-02: 0 commits2026-03-03: 0 commits2026-03-04: 0 commits2026-03-05: 0 commits2026-03-06: 2 commits2026-03-07: 0 commits2026-03-08: 0 commits2026-03-09: 0 commits2026-03-10: 0 commits2026-03-11: 1 commit2026-03-12: 0 commits2026-03-13: 1 commit2026-03-14: 0 commits2026-03-15: 2 commits2026-03-16: 0 commits2026-03-17: 0 commits2026-03-18: 0 commits2026-03-19: 0 commits2026-03-20: 0 commits2026-03-21: 0 commits2026-03-22: 0 commits2026-03-23: 0 commits2026-03-24: 1 commit2026-03-25: 0 commits2026-03-26: 0 commits2026-03-27: 0 commits2026-03-28: 1 commit2026-03-29: 0 commits2026-03-30: 0 commits2026-03-31: 0 commits2026-04-01: 1 commit2026-04-02: 0 commits2026-04-03: 0 commits2026-04-04: 8 commits2026-04-05: 0 commits2026-04-06: 0 commits2026-04-07: 1 commit2026-04-08: 1 commit2026-04-09: 0 commits2026-04-10: 0 commits2026-04-11: 1 commit2026-04-12: 2 commits2026-04-13: 0 commits2026-04-14: 1 commit2026-04-15: 0 commits2026-04-16: 0 commits2026-04-17: 0 commits2026-04-18: 1 commit2026-04-19: 0 commits2026-04-20: 0 commits2026-04-21: 1 commit2026-04-22: 0 commits2026-04-23: 2 commits2026-04-24: 0 commits2026-04-25: 1 commit2026-04-26: 0 commits2026-04-27: 0 commits2026-04-28: 0 commits2026-04-29: 0 commits2026-04-30: 0 commits2026-05-01: 0 commits2026-05-02: 2 commits2026-05-03: 3 commits2026-05-04: 1 commit2026-05-05: 1 commit2026-05-06: 0 commits2026-05-07: 1 commit2026-05-08: 0 commits2026-05-09: 2 commits2026-05-10: 0 commits2026-05-11: 0 commits2026-05-12: 0 commits2026-05-13: 0 commits2026-05-14: 0 commits2026-05-15: 0 commits2026-05-16: 1 commit2026-05-17: 0 commits2026-05-18: 0 commits2026-05-19: 0 commits2026-05-20: 0 commits2026-05-21: 0 commits2026-05-22: 3 commits2026-05-23: 0 commits2026-05-24: 0 commits2026-05-25: 0 commits2026-05-26: 0 commits2026-05-27: 0 commits2026-05-28: 1 commit2026-05-29: 0 commits2026-05-30: 0 commits2026-05-31: 3 commits2026-06-01: 0 commits2026-06-02: 0 commits2026-06-03: 1 commit2026-06-04: 2 commits2026-06-05: 2 commits2026-06-06: 1 commit2026-06-07: 2 commits2026-06-08: 0 commits2026-06-09: 1 commit2026-06-10: 0 commits2026-06-11: 2 commits2026-06-12: 0 commits2026-06-13: 1 commit2026-06-14: 1 commit2026-06-15: 0 commits2026-06-16: 0 commits2026-06-17: 2 commits2026-06-18: 2 commits2026-06-19: 0 commits2026-06-20: 0 commits2026-06-21: 0 commits2026-06-22: 0 commits2026-06-23: 0 commits2026-06-24: 0 commits2026-06-25: 0 commits2026-06-26: 0 commits2026-06-27: 2 commits2026-06-28: 0 commits2026-06-29: 0 commits2026-06-30: 0 commits2026-07-01: 1 commit2026-07-02: 0 commits2026-07-03: 0 commits2026-07-04: 0 commits2026-07-05: 0 commits2026-07-06: 0 commits2026-07-07: 1 commit2026-07-08: 0 commits2026-07-09: 0 commits2026-07-10: 0 commits2026-07-11: 0 commits2026-07-12: 0 commits2026-07-13: 0 commits2026-07-14: 0 commits2026-07-15: 1 commit2026-07-16: 1 commit2026-07-17: 1 commit2026-07-18: 0 commits2026-07-19: 0 commits2026-07-20: 0 commits2026-07-21: 0 commits2026-07-22: 2 commits2026-07-23: 0 commits2026-07-24: 1 commit2026-07-25: 0 commits2026-07-26: 0 commits2026-07-27: 0 commits2026-07-28: 2 commits2026-07-29: 0 commits2026-07-30: 0 commits2026-07-31: 0 commits2026-08-01: 0 commits
191 commits in the last yearLessMore

Signals and awards

derived from tracked data
  • Widely adopted

    27,180 stars

  • Actively maintained

    Pushed within 48 hours

  • Continuous integration

    Automated checks passing

  • Repeat trending

    7 trending appearances

What heretic does

Heretic is a specialized tool that automates the process of 'abliterating' or decensoring large language models. It uses Tree-structured Parzen Estimator (TPE) parameter optimization, powered by Optuna, to find the optimal configuration that minimizes refusal rates while preserving the model's original intelligence and capabilities. By co-minimizing refusals and KL divergence, it provides a fully automatic way to fine-tune models without requiring deep knowledge of transformer architectures. It supports various dense models and some Mixture of Experts (MoE) architectures.

Heretic is designed for AI researchers, developers, and local LLM enthusiasts who want to remove safety filters from open-source models. It requires technical familiarity with running command-line tools and managing LLM environments.

  • Automated Optimization: Uses Optuna-backed TPE to automatically find the best parameters for decensoring.
  • KL Divergence Minimization: Ensures the modified model retains the intelligence and characteristics of the original.
  • Broad Compatibility: Supports most dense models, multimodal models, and select MoE architectures.
  • User-Friendly CLI: Designed to be run easily by anyone familiar with basic command-line interfaces.
  • Metrics Tracking: Provides detailed reporting on refusal rates and model degradation during the process.

Where teams use it

AI Researchers

Study the impact of alignment techniques and explore the boundary conditions of model safety mechanisms.

Local Model Enthusiasts

Remove restrictive guardrails from open-source models for uncensored creative writing or roleplay.

Red Teamers

Analyze how easily safety alignments can be bypassed to improve future model defenses.

Developers

Fine-tune models for specific applications where default safety filters produce unacceptable false positives.

Getting started: pip install heretic-llm

README

master branch
Logo

Heretic: Fully automatic censorship removal for language models

Discord Matrix Follow us on Hugging Face Codeberg mirror

#1 Repository of the Day

Heretic is a tool that removes censorship (aka "safety alignment") from transformer-based language models without expensive post-training. It combines an advanced implementation of directional ablation, also known as "abliteration" (Arditi et al. 2024, Lai 2025 (1, 2)), with a TPE-based parameter optimizer powered by Optuna.

This approach enables Heretic to work completely automatically. Heretic finds high-quality abliteration parameters by co-minimizing the number of refusals and the KL divergence from the original model. This results in a decensored model that retains as much of the original model's intelligence as possible. Using Heretic does not require an understanding of transformer internals. In fact, anyone who knows how to run a command-line program can use Heretic to decensor language models.

Heretic supports most dense models, including many multimodal models, several different MoE architectures, and even some hybrid models like Qwen3.5. Pure state-space models and certain other research architectures are not yet supported out of the box.

Screenshot

 

Running unsupervised with the default configuration, Heretic can produce decensored models that rival the quality of abliterations created manually by human experts:

Model Refusals for "harmful" prompts KL divergence from original model for "harmless" prompts
google/gemma-3-12b-it (original) 97/100 0 (by definition)
mlabonne/gemma-3-12b-it-abliterated-v2 3/100 1.04
huihui-ai/gemma-3-12b-it-abliterated 3/100 0.45
p-e-w/gemma-3-12b-it-heretic (ours) 3/100 0.16

The Heretic version, generated without any human effort, achieves the same level of refusal suppression as other abliterations, but at a much lower KL divergence, indicating less damage to the original model's capabilities. (You can reproduce those numbers using Heretic's built-in evaluation functionality, e.g. heretic --model google/gemma-3-12b-it --evaluate-model p-e-w/gemma-3-12b-it-heretic. Note that the exact values might be platform- and hardware-dependent. The table above was compiled using PyTorch 2.8 on an RTX 5090.)

Of course, mathematical metrics and automated benchmarks never tell the whole story, and are no substitute for human evaluation. Models generated with Heretic have been well-received by users (links and emphasis added):

"I was skeptical before, but I just downloaded GPT-OSS 20B Heretic model and holy shit. It gives properly formatted long responses to sensitive topics, using the exact uncensored words that you would expect from an uncensored model, produces markdown format tables with details and whatnot. Looks like this is the best abliterated version of this model so far..." (Link to comment)

"Heretic GPT 20b seems to be the best uncensored model I have tried yet. It doesn't destroy a the model's intelligence and it is answering prompts normally would be rejected by the base model." (Link to comment)

"[Qwen3-4B-Instruct-2507-heretic] Has been the best unquantized abliterated model that I have been able to run on 16gb vram." (Link to comment)

Heretic models have also been independently benchmarked using standard metrics like MMLU and GSM8K, and have been found to compare favorably with models produced by competing abliteration tools: 1, 2.

The community has created and published well over 4000 models with Heretic.

Usage

Prepare a Python 3.10+ environment with PyTorch 2.2+ installed as appropriate for your hardware. Then run:

pip install -U heretic-llm
heretic Qwen/Qwen3-4B-Instruct-2507

Replace Qwen/Qwen3-4B-Instruct-2507 with whatever model you want to decensor.

Important

While PyTorch 2.2 is the minimum version of PyTorch needed for Heretic to work, some models and configurations might require features only found in later versions. For example, loading MXFP4-quantized models like gpt-oss uses torch.accelerator, which was added in PyTorch 2.6.

Tip

Heretic uses uv for dependency management, and the repository includes a uv.lock file pinning every package version. If you already use uv (and you probably should!), you can just clone the repo and run Heretic with uv run heretic, which ensures that your dependencies match those used by the developers, improving reliability and security.

The process is fully automatic and does not require configuration; however, Heretic has a variety of configuration parameters that can be changed for greater control. Run heretic --help to see available command-line options, or look at config.default.toml if you prefer to use a configuration file.

At the start of a program run, Heretic benchmarks the system to determine the optimal batch size to make the most of the available hardware. On an RTX 3090, with the default configuration, decensoring Qwen3-4B-Instruct-2507 takes about 20-30 minutes. Note that Heretic supports model quantization with bitsandbytes, which can drastically reduce the amount of VRAM required to process models. Set the quantization option to bnb_4bit to enable quantization.

After Heretic has finished decensoring a model, you are given the option to save the model, upload it to Hugging Face, chat with it to test how well it works, run standard benchmarks on it, or any combination of those actions.

Research features

In addition to its primary function of removing model censorship, Heretic also provides features designed to support research into the semantics of model internals (interpretability). To use those features, you need to install Heretic with the optional research extra:

pip install -U 'heretic-llm[research]'

This gives you access to the following functionality:

Generate plots of residual vectors by passing --plot-residuals

When run with this flag, Heretic will:

  1. Compute residual vectors (hidden states) for the first output token, for each transformer layer, for both "harmful" and "harmless" prompts.
  2. Perform a PaCMAP projection from residual space to 2D-space.
  3. Left-right align the projections of "harmful"/"harmless" residuals by their geometric medians to make projections for consecutive layers more similar. Additionally, PaCMAP is initialized with the previous layer's projections for each new layer, minimizing disruptive transitions.
  4. Scatter-plot the projections, generating a PNG image for each layer.
  5. Generate an animation showing how residuals transform between layers, as an animated GIF.
Plot of residual vectors

See the configuration file for options that allow you to control various aspects of the generated plots.

Note that PaCMAP is an expensive operation that is performed on the CPU. For larger models, it can take an hour or more to compute projections for all layers.

Print details about residual geometry by passing --print-residual-geometry

If you are interested in a quantitative analysis of how residual vectors for "harmful" and "harmless" prompts relate to each other, this flag gives you the following table, packed with metrics that can facilitate understanding the same (for gemma-3-270m-it in this case):

┏━━━━━━━┳━━━━━━━━┳━━━━━━━━━━┳━━━━━━━━━┳━━━━━━━━━━┳━━━━━━━━━┳━━━━━━━━━━┳━━━━━━━━━━┳━━━━━━━━━━┳━━━━━━━━━━┳━━━━━━━━━━┳━━━━━━━━━┳━━━━━━━━━┳━━━━━━━━┓
┃ Layer ┃ S(g,b) ┃ S(g*,b*) ┃  S(g,r) ┃ S(g*,r*) ┃  S(b,r) ┃ S(b*,r*) ┃      |g| ┃     |g*| ┃      |b| ┃     |b*| ┃     |r| ┃    |r*| ┃   Silh ┃
┡━━━━━━━╇━━━━━━━━╇━━━━━━━━━━╇━━━━━━━━━╇━━━━━━━━━━╇━━━━━━━━━╇━━━━━━━━━━╇━━━━━━━━━━╇━━━━━━━━━━╇━━━━━━━━━━╇━━━━━━━━━━╇━━━━━━━━━╇━━━━━━━━━╇━━━━━━━━┩
│     1 │ 1.0000 │   1.0000 │ -0.4311 │  -0.4906 │ -0.4254 │  -0.4847 │   170.29 │   170.49 │   169.78 │   169.85 │    1.19 │    1.31 │ 0.0480 │
│     2 │ 1.0000 │   1.0000 │  0.4297 │   0.4465 │  0.4365 │   0.4524 │   768.55 │   768.77 │   771.32 │   771.36 │    6.39 │    5.76 │ 0.0745 │
│     3 │ 0.9999 │   1.0000 │ -0.5699 │  -0.5577 │ -0.5614 │  -0.5498 │  1020.98 │  1021.13 │  1013.80 │  1014.71 │   12.70 │   11.60 │ 0.0920 │
│     4 │ 0.9999 │   1.0000 │  0.6582 │   0.6553 │  0.6659 │   0.6627 │  1356.39 │  1356.20 │  1368.71 │  1367.95 │   18.62 │   17.84 │ 0.0957 │
│     5 │ 0.9987 │   0.9990 │ -0.6880 │  -0.6761 │ -0.6497 │  -0.6418 │   766.54 │   762.25 │   731.75 │   732.42 │   51.97 │   45.24 │ 0.1018 │
│     6 │ 0.9998 │   0.9998 │ -0.1983 │  -0.2312 │ -0.1811 │  -0.2141 │  2417.35 │  2421.08 │  2409.18 │  2411.40 │   43.06 │   43.47 │ 0.0900 │
│     7 │ 0.9998 │   0.9997 │ -0.5258 │  -0.5746 │ -0.5072 │  -0.5560 │  3444.92 │  3474.99 │  3400.01 │  3421.63 │   86.94 │   94.38 │ 0.0492 │
│     8 │ 0.9990 │   0.9991 │  0.8235 │   0.8312 │  0.8479 │   0.8542 │  4596.54 │  4615.62 │  4918.32 │  4934.20 │  384.87 │  377.87 │ 0.2278 │
│     9 │ 0.9992 │   0.9992 │  0.5335 │   0.5441 │  0.5678 │   0.5780 │  5322.30 │  5316.96 │  5468.65 │  5466.98 │  265.68 │  267.28 │ 0.1318 │
│    10 │ 0.9974 │   0.9973 │  0.8189 │   0.8250 │  0.8579 │   0.8644 │  5328.81 │  5325.63 │  5953.35 │  5985.15 │  743.95 │  779.74 │ 0.2863 │
│    11 │ 0.9977 │   0.9978 │  0.4262 │   0.4045 │  0.4862 │   0.4645 │  9644.02 │  9674.06 │  9983.47 │  9990.28 │  743.28 │  726.99 │ 0.1576 │
│    12 │ 0.9904 │   0.9907 │  0.4384 │   0.4077 │  0.5586 │   0.5283 │ 10257.40 │ 10368.50 │ 11114.51 │ 11151.21 │ 1711.18 │ 1664.69 │ 0.1890 │
│    13 │ 0.9867 │   0.9874 │  0.4007 │   0.3680 │  0.5444 │   0.5103 │ 12305.12 │ 12423.75 │ 13440.31 │ 13432.47 │ 2386.43 │ 2282.47 │ 0.1293 │
│    14 │ 0.9921 │   0.9922 │  0.3198 │   0.2682 │  0.4364 │   0.3859 │ 16929.16 │ 17080.37 │ 17826.97 │ 17836.03 │ 2365.23 │ 2301.87 │ 0.1282 │
│    15 │ 0.9846 │   0.9850 │  0.1198 │   0.0963 │  0.2913 │   0.2663 │ 16858.58 │ 16949.44 │ 17496.00 │ 17502.88 │ 3077.08 │ 3029.60 │ 0.1611 │
│    16 │ 0.9686 │   0.9689 │ -0.0029 │  -0.0254 │  0.2457 │   0.2226 │ 18912.77 │ 19074.86 │ 19510.56 │ 19559.62 │ 4848.35 │ 4839.75 │ 0.1516 │
│    17 │ 0.9782 │   0.9784 │ -0.0174 │  -0.0381 │  0.1908 │   0.1694 │ 27098.09 │ 27273.00 │ 27601.12 │ 27653.12 │ 5738.19 │ 5724.21 │ 0.1641 │
│    18 │ 0.9184 │   0.9196 │  0.1343 │   0.1430 │  0.5155 │   0.5204 │   190.16 │   190.35 │   219.91 │   220.62 │   87.82 │   87.59 │ 0.1855 │
└───────┴────────┴──────────┴─────────┴──────────┴─────────┴──────────┴──────────┴──────────┴──────────┴──────────┴─────────┴─────────┴────────┘
g = mean of residual vectors for good prompts
g* = geometric median of residual vectors for good prompts
b = mean of residual vectors for bad prompts
b* = geometric median of residual vectors for bad prompts
r = residual direction for means (i.e., b - g)
r* = residual direction for geometric medians (i.e., b* - g*)
S(x,y) = cosine similarity of x and y
|x| = L2 norm of x
Silh = Mean silhouette coefficient of residuals for good/bad clusters

How Heretic works

Heretic implements a parametrized variant of directional ablation. For each supported transformer component (currently, attention out-projection and MLP down-projection), it identifies the associated matrices in each transformer layer, and orthogonalizes them with respect to the relevant "residual direction", inhibiting the expression of that direction in the result of multiplications with that matrix.

Residual directions are computed for each layer as a difference-of-means between the first-token residuals for "harmful" and "harmless" example prompts.

The ablation process is controlled by several optimizable parameters:

  • direction_index: Either the index of a residual direction, or the special value per layer, indicating that each layer should be ablated using the residual direction associated with that layer.
  • max_weight, max_weight_position, min_weight, and min_weight_distance: For each component, these parameters describe the shape and position of the ablation weight kernel over the layers. The following diagram illustrates this:
Explanation

 

Heretic's main innovations over existing abliteration systems are:

  • The shape of the ablation weight kernel is highly flexible, which, combined with automatic parameter optimization, can improve the compliance/quality tradeoff. Non-constant ablation weights were previously explored by Maxime Labonne in gemma-3-12b-it-abliterated-v2.
  • The residual direction index is a float rather than an integer. For non-integral values, the two nearest residual direction vectors are linearly interpolated. This unlocks a vast space of additional directions beyond the ones identified by the difference-of-means computation, and often enables the optimization process to find a better direction than that belonging to any individual layer.
  • Ablation parameters are chosen separately for each component. I have found that MLP interventions tend to be more damaging to the model than attention interventions, so using different ablation weights can squeeze out some extra performance.

Prior art

I'm aware of the following publicly available implementations of abliteration techniques:

Note that Heretic was written from scratch, and does not reuse code from any of those projects.

Acknowledgments

The development of Heretic was informed by:

Citation

If you use Heretic for your research, please cite it using the following BibTeX entry:

@misc{heretic,
  author = {Weidmann, Philipp Emanuel},
  title = {Heretic: Fully automatic censorship removal for language models},
  year = {2025},
  publisher = {GitHub},
  journal = {GitHub repository},
  howpublished = {\url{https://github.com/p-e-w/heretic}}
}

License

Copyright © 2025-2026 Philipp Emanuel Weidmann (pew@worldwidemann.com) + contributors

This program is free software: you can redistribute it and/or modify it under the terms of the GNU Affero General Public License as published by the Free Software Foundation, either version 3 of the License, or (at your option) any later version.

This program is distributed in the hope that it will be useful, but WITHOUT ANY WARRANTY; without even the implied warranty of MERCHANTABILITY or FITNESS FOR A PARTICULAR PURPOSE. See the GNU Affero General Public License for more details.

You should have received a copy of the GNU Affero General Public License along with this program. If not, see https://www.gnu.org/licenses/.

By contributing to this project, you agree to release your contributions under the same license.

View on GitHub

Recent activity

commits and pull requests

Releases and announcements

5 total
  1. v1.4.0v1.4.0Jun 14, 20261.3K downloads

    ## Changes * @p-e-w implemented automatically reproducing a model from a `reproduce.json` file in https://github.com/p-e-w/heretic/pull/326. @Vinay-Umrethe fixed issues in that implementation in https://github.com/p-e-w/heretic/pull/352. * @rocker-zhang added support for plain text files as prompt datasets in https://github.com/p-e-w/heretic/pull/337, based on earlier work by @ricyoung in https://github.com/p-e-w/heretic/pull/103. * @anrp fixed LoRA export in https://github.com/p-e-w/heretic/pull/321 * @zaakirio implemented saving the processor for multimodal models in https://github.com/p-e-w/heretic/pull/353 * @MoonRide303 added support for gemma-4-12B-it in https://github.com/p-e-w/heretic/pull/350 * @coder3101 added support for LiquidAI/LFM2.5 in https://github.com/p-e-w/heretic/pull/344 * @UnstableLlama add a config file for suppressing humor in https://github.com/p-e-w/heretic/pull/340 * @umran666 fixed a trial counting bug in https://github.com/p-e-w/heretic/pull/357 * @umran666 made resetting the model null-safe to handle study cancellations in https://github.com/p-e-w/heretic/pull/367 * @umran666 improved exception formatting in https://github.com/p-e-w/heretic

  2. v1.3.0v1.3.0May 5, 2026324 downloads

    ## Changes * @Vinay-Umrethe (who had previously contributed under the username @Vinayyyy7) implemented reproducible runs in https://github.com/p-e-w/heretic/pull/191. @p-e-w revised and improved that implementation in https://github.com/p-e-w/heretic/pull/303. * @magiccodingman reduced peak VRAM usage in https://github.com/p-e-w/heretic/pull/239. @olekssy fixed a bug in that implementation in https://github.com/p-e-w/heretic/pull/301. * @farolone added support for Qwen3.5 models in https://github.com/p-e-w/heretic/pull/187 * @MoonRide303 added support for Gemma 4 models in https://github.com/p-e-w/heretic/pull/287 * @erm14254 made sure all abliterable components across layers are displayed in https://github.com/p-e-w/heretic/pull/215 * @cpagac fixed VRAM usage reporting for multi-GPU setups in https://github.com/p-e-w/heretic/pull/169 * @cpagac fixed a division-by-zero error in the evaluator in https://github.com/p-e-w/heretic/pull/225 * @spikymoth improved automatic response prefix determination with a two-step process in https://github.com/p-e-w/heretic/pull/194 * @spikymoth added model card generation for local models with an existing README in https://github.com/p-e-

  3. v1.2.0v1.2.0Feb 14, 202627 downloads

    ## Changes * @noctrex added a `max_memory` setting to limit memory usage in https://github.com/p-e-w/heretic/pull/83 * @spikymoth added a mechanism to avoid excessive low-divergence iteration in https://github.com/p-e-w/heretic/pull/73 * @accemlcc implemented a new LoRA-based abliteration engine with support for 4-bit quantization in https://github.com/p-e-w/heretic/pull/60 * @accemlcc added enumeration of all available GPUs on startup in https://github.com/p-e-w/heretic/pull/86 * @Vinayyyy7 added the ability to run more trials after optimization is complete in https://github.com/p-e-w/heretic/pull/76 * @anrp fixed MXFP4 loading in https://github.com/p-e-w/heretic/pull/107 * @anrp refactored the save machinery in https://github.com/p-e-w/heretic/pull/110 * @anrp added broad support for VL models in https://github.com/p-e-w/heretic/pull/108 * @anrp implemented saving and resuming optimization progress in https://github.com/p-e-w/heretic/pull/106, https://github.com/p-e-w/heretic/pull/119, and https://github.com/p-e-w/heretic/pull/116 * @spikymoth implemented Magnitude-Preserving Orthogonal Ablation in https://github.com/p-e-w/heretic/pull/52 * @salmanmkc upgraded GitHub A

  4. v1.1.0v1.1.0Dec 10, 202527 downloads

    ## Changes * @mbarnson added basic MPS (Apple Silicon) support in https://github.com/p-e-w/heretic/pull/5 * @red40maxxer reduced memory usage in https://github.com/p-e-w/heretic/pull/15 * @Ooooze added IBM Granite MoE support in https://github.com/p-e-w/heretic/pull/14 * @kldzj added multi-GPU support in https://github.com/p-e-w/heretic/pull/17 and https://github.com/p-e-w/heretic/pull/32 * @ricyoung fixed an error when Hugging Face user profile fields are missing in https://github.com/p-e-w/heretic/pull/20 * @tymat added support for MXFP4 quantized models with Triton tensors in https://github.com/p-e-w/heretic/pull/28 * @spikymoth improved support for loading local datasets in https://github.com/p-e-w/heretic/pull/33 * @kldzj added support for models that require `trust_remote_code` in https://github.com/p-e-w/heretic/pull/31 * @Vinayyyy7 added notebook (Colab/Kaggle) compatibility in https://github.com/p-e-w/heretic/pull/42 * @Vinayyyy7 fixed loading for certain models that default to the float32 dtype in https://github.com/p-e-w/heretic/pull/44 * @spikymoth improved refusal detection in https://github.com/p-e-w/heretic/pull/45 * @red40maxxer added a PR title lint t

  5. v1.0.1v1.0.1Nov 16, 202527 downloads

    First public release

Code frequency

additions and deletions
+3.8K-3.8KWeek of 2025-09-21: +3,823 linesWeek of 2025-09-21: -70 linesWeek of 2025-09-28: +59 linesWeek of 2025-09-28: -2 linesWeek of 2025-10-05: +50 linesWeek of 2025-10-05: -25 linesWeek of 2025-10-12: +100 linesWeek of 2025-10-12: -11 linesWeek of 2025-10-19: +202 linesWeek of 2025-10-19: -127 linesWeek of 2025-10-26: +149 linesWeek of 2025-10-26: -50 linesWeek of 2025-11-02: +121 linesWeek of 2025-11-02: -18 linesWeek of 2025-11-09: +227 linesWeek of 2025-11-09: -210 linesWeek of 2025-11-16: +408 linesWeek of 2025-11-16: -34 linesWeek of 2025-11-23: +203 linesWeek of 2025-11-23: -23 linesWeek of 2025-11-30: +1,284 linesWeek of 2025-11-30: -64 linesWeek of 2025-12-07: +261 linesWeek of 2025-12-07: -89 linesWeek of 2025-12-14: +2,704 linesWeek of 2025-12-14: -1,733 linesWeek of 2025-12-21: +228 linesWeek of 2025-12-21: -94 linesWeek of 2025-12-28: +105 linesWeek of 2025-12-28: -41 linesWeek of 2026-01-04: +9 linesWeek of 2026-01-04: -0 linesWeek of 2026-01-11: +182 linesWeek of 2026-01-11: -2 linesWeek of 2026-01-18: +192 linesWeek of 2026-01-18: -60 linesWeek of 2026-01-25: +5 linesWeek of 2026-01-25: -3 linesWeek of 2026-02-01: +172 linesWeek of 2026-02-01: -23 linesWeek of 2026-02-08: +293 linesWeek of 2026-02-08: -242 linesWeek of 2026-02-15: +40 linesWeek of 2026-02-15: -10 linesWeek of 2026-02-22: +29 linesWeek of 2026-02-22: -13 linesWeek of 2026-03-01: +39 linesWeek of 2026-03-01: -19 linesWeek of 2026-03-08: +11 linesWeek of 2026-03-08: -5 linesWeek of 2026-03-15: +160 linesWeek of 2026-03-15: -104 linesWeek of 2026-03-22: +791 linesWeek of 2026-03-22: -16 linesWeek of 2026-03-29: +218 linesWeek of 2026-03-29: -221 linesWeek of 2026-04-05: +1,047 linesWeek of 2026-04-05: -142 linesWeek of 2026-04-12: +208 linesWeek of 2026-04-12: -121 linesWeek of 2026-04-19: +463 linesWeek of 2026-04-19: -331 linesWeek of 2026-04-26: +35 linesWeek of 2026-04-26: -32 linesWeek of 2026-05-03: +282 linesWeek of 2026-05-03: -108 linesWeek of 2026-05-10: +8 linesWeek of 2026-05-10: -7 linesWeek of 2026-05-17: +9 linesWeek of 2026-05-17: -8 linesWeek of 2026-05-24: +2 linesWeek of 2026-05-24: -0 linesWeek of 2026-05-31: +235 linesWeek of 2026-05-31: -134 linesWeek of 2026-06-07: +782 linesWeek of 2026-06-07: -245 linesWeek of 2026-06-14: +155 linesWeek of 2026-06-14: -112 linesWeek of 2026-06-21: +717 linesWeek of 2026-06-21: -274 linesWeek of 2026-06-28: +214 linesWeek of 2026-06-28: -14 linesWeek of 2026-07-05: +1,177 linesWeek of 2026-07-05: -389 linesWeek of 2026-07-12: +90 linesWeek of 2026-07-12: -27 linesWeek of 2026-07-19: +277 linesWeek of 2026-07-19: -145 linesWeek of 2026-07-26: +93 linesWeek of 2026-07-26: -97 linesWeek of 2026-08-02: +151 linesWeek of 2026-08-02: -151 linesSep 21, 2025Aug 2, 2026
+18K lines added, -5.6K removed over the last year.

Commits per week

last 52 weeks
200Week of 2025-08-03: 0 commitsWeek of 2025-08-10: 0 commitsWeek of 2025-08-17: 0 commitsWeek of 2025-08-24: 0 commitsWeek of 2025-08-31: 0 commitsWeek of 2025-09-07: 0 commitsWeek of 2025-09-14: 0 commitsWeek of 2025-09-21: 10 commitsWeek of 2025-09-28: 1 commitsWeek of 2025-10-05: 2 commitsWeek of 2025-10-12: 6 commitsWeek of 2025-10-19: 9 commitsWeek of 2025-10-26: 4 commitsWeek of 2025-11-02: 2 commitsWeek of 2025-11-09: 2 commitsWeek of 2025-11-16: 20 commitsWeek of 2025-11-23: 5 commitsWeek of 2025-11-30: 5 commitsWeek of 2025-12-07: 10 commitsWeek of 2025-12-14: 5 commitsWeek of 2025-12-21: 5 commitsWeek of 2025-12-28: 3 commitsWeek of 2026-01-04: 1 commitsWeek of 2026-01-11: 2 commitsWeek of 2026-01-18: 3 commitsWeek of 2026-01-25: 2 commitsWeek of 2026-02-01: 2 commitsWeek of 2026-02-08: 13 commitsWeek of 2026-02-15: 3 commitsWeek of 2026-02-22: 1 commitsWeek of 2026-03-01: 2 commitsWeek of 2026-03-08: 2 commitsWeek of 2026-03-15: 2 commitsWeek of 2026-03-22: 2 commitsWeek of 2026-03-29: 9 commitsWeek of 2026-04-05: 3 commitsWeek of 2026-04-12: 4 commitsWeek of 2026-04-19: 4 commitsWeek of 2026-04-26: 2 commitsWeek of 2026-05-03: 8 commitsWeek of 2026-05-10: 1 commitsWeek of 2026-05-17: 3 commitsWeek of 2026-05-24: 1 commitsWeek of 2026-05-31: 9 commitsWeek of 2026-06-07: 6 commitsWeek of 2026-06-14: 5 commitsWeek of 2026-06-21: 2 commitsWeek of 2026-06-28: 1 commitsWeek of 2026-07-05: 1 commitsWeek of 2026-07-12: 3 commitsWeek of 2026-07-19: 3 commitsWeek of 2026-07-26: 2 commitsAug 3, 2025Jul 26, 2026
191 commits in the last 52 weeks.

When work happens

weekday and hour
SunMonTueWedThuFriSat036912151821Sun 0:00 — 0 commitsSun 1:00 — 0 commitsSun 2:00 — 1 commitsSun 3:00 — 0 commitsSun 4:00 — 0 commitsSun 5:00 — 0 commitsSun 6:00 — 2 commitsSun 7:00 — 1 commitsSun 8:00 — 2 commitsSun 9:00 — 7 commitsSun 10:00 — 3 commitsSun 11:00 — 5 commitsSun 12:00 — 1 commitsSun 13:00 — 2 commitsSun 14:00 — 2 commitsSun 15:00 — 2 commitsSun 16:00 — 4 commitsSun 17:00 — 4 commitsSun 18:00 — 1 commitsSun 19:00 — 1 commitsSun 20:00 — 0 commitsSun 21:00 — 0 commitsSun 22:00 — 0 commitsSun 23:00 — 0 commitsMon 0:00 — 0 commitsMon 1:00 — 1 commitsMon 2:00 — 0 commitsMon 3:00 — 0 commitsMon 4:00 — 0 commitsMon 5:00 — 1 commitsMon 6:00 — 0 commitsMon 7:00 — 0 commitsMon 8:00 — 0 commitsMon 9:00 — 0 commitsMon 10:00 — 3 commitsMon 11:00 — 2 commitsMon 12:00 — 1 commitsMon 13:00 — 0 commitsMon 14:00 — 0 commitsMon 15:00 — 1 commitsMon 16:00 — 0 commitsMon 17:00 — 0 commitsMon 18:00 — 0 commitsMon 19:00 — 1 commitsMon 20:00 — 0 commitsMon 21:00 — 2 commitsMon 22:00 — 1 commitsMon 23:00 — 0 commitsTue 0:00 — 0 commitsTue 1:00 — 1 commitsTue 2:00 — 0 commitsTue 3:00 — 0 commitsTue 4:00 — 0 commitsTue 5:00 — 2 commitsTue 6:00 — 1 commitsTue 7:00 — 1 commitsTue 8:00 — 3 commitsTue 9:00 — 0 commitsTue 10:00 — 1 commitsTue 11:00 — 2 commitsTue 12:00 — 1 commitsTue 13:00 — 5 commitsTue 14:00 — 0 commitsTue 15:00 — 1 commitsTue 16:00 — 3 commitsTue 17:00 — 0 commitsTue 18:00 — 2 commitsTue 19:00 — 2 commitsTue 20:00 — 0 commitsTue 21:00 — 0 commitsTue 22:00 — 1 commitsTue 23:00 — 0 commitsWed 0:00 — 1 commitsWed 1:00 — 0 commitsWed 2:00 — 0 commitsWed 3:00 — 0 commitsWed 4:00 — 0 commitsWed 5:00 — 0 commitsWed 6:00 — 2 commitsWed 7:00 — 0 commitsWed 8:00 — 0 commitsWed 9:00 — 4 commitsWed 10:00 — 4 commitsWed 11:00 — 3 commitsWed 12:00 — 4 commitsWed 13:00 — 1 commitsWed 14:00 — 3 commitsWed 15:00 — 1 commitsWed 16:00 — 4 commitsWed 17:00 — 2 commitsWed 18:00 — 1 commitsWed 19:00 — 0 commitsWed 20:00 — 0 commitsWed 21:00 — 0 commitsWed 22:00 — 1 commitsWed 23:00 — 0 commitsThu 0:00 — 0 commitsThu 1:00 — 1 commitsThu 2:00 — 0 commitsThu 3:00 — 0 commitsThu 4:00 — 0 commitsThu 5:00 — 0 commitsThu 6:00 — 0 commitsThu 7:00 — 0 commitsThu 8:00 — 0 commitsThu 9:00 — 0 commitsThu 10:00 — 0 commitsThu 11:00 — 2 commitsThu 12:00 — 3 commitsThu 13:00 — 0 commitsThu 14:00 — 3 commitsThu 15:00 — 2 commitsThu 16:00 — 1 commitsThu 17:00 — 1 commitsThu 18:00 — 3 commitsThu 19:00 — 1 commitsThu 20:00 — 0 commitsThu 21:00 — 0 commitsThu 22:00 — 0 commitsThu 23:00 — 0 commitsFri 0:00 — 1 commitsFri 1:00 — 0 commitsFri 2:00 — 1 commitsFri 3:00 — 1 commitsFri 4:00 — 1 commitsFri 5:00 — 0 commitsFri 6:00 — 0 commitsFri 7:00 — 0 commitsFri 8:00 — 1 commitsFri 9:00 — 0 commitsFri 10:00 — 1 commitsFri 11:00 — 1 commitsFri 12:00 — 2 commitsFri 13:00 — 5 commitsFri 14:00 — 5 commitsFri 15:00 — 2 commitsFri 16:00 — 2 commitsFri 17:00 — 0 commitsFri 18:00 — 1 commitsFri 19:00 — 0 commitsFri 20:00 — 2 commitsFri 21:00 — 0 commitsFri 22:00 — 0 commitsFri 23:00 — 1 commitsSat 0:00 — 1 commitsSat 1:00 — 0 commitsSat 2:00 — 0 commitsSat 3:00 — 0 commitsSat 4:00 — 0 commitsSat 5:00 — 0 commitsSat 6:00 — 2 commitsSat 7:00 — 1 commitsSat 8:00 — 9 commitsSat 9:00 — 5 commitsSat 10:00 — 2 commitsSat 11:00 — 2 commitsSat 12:00 — 0 commitsSat 13:00 — 4 commitsSat 14:00 — 1 commitsSat 15:00 — 1 commitsSat 16:00 — 3 commitsSat 17:00 — 2 commitsSat 18:00 — 4 commitsSat 19:00 — 3 commitsSat 20:00 — 0 commitsSat 21:00 — 0 commitsSat 22:00 — 0 commitsSat 23:00 — 0 commits
Commit volume by weekday and hour (UTC). Larger dots mean more commits.

Who is committing

last 52 weeks
Maintainer commits107 (56%)
Community commits85 (44%)

192 commits in total over the last year.

DateListRankStars gained
Mar 15, 2026daily#12+292
Mar 14, 2026daily#19+212
Mar 13, 2026daily#18+189
Feb 19, 2026daily#21+127
Feb 18, 2026daily#12+243
Feb 17, 2026daily#12+195
Feb 16, 2026daily#6+305
  • public-apis/public-apis

    A collective list of free APIs

    454.9K stars · Python

  • donnemartin/system-design-primer

    Learn how to design large-scale systems. Prep for the system design interview. Includes Anki flashcards.

    362.2K stars · Python

  • practical-tutorials/project-based-learning

    Curated list of project-based tutorials

    277.2K stars · Python

  • affaan-m/ECC

    The agent harness performance optimization system. Skills, instincts, memory, security, and research-first development for Claude Code, Codex, Opencode, Cursor and beyond.

    238.5K stars · JavaScript

  • affaan-m/ECC

    The agent harness performance optimization system. Skills, instincts, memory, security, and research-first development for Claude Code, Codex, Opencode, Cursor and beyond.

    234.7K stars · JavaScript

  • NousResearch/hermes-agent

    The agent that grows with you

    227K stars · Python