A deliberative governance engine that decides whether, how, and under what constraints an LLM should respond β before a single token is generated.
Quickstart Β· How it works Β· Benchmarks Β· Configuration Β· Docs Β· Contributing
Most AI safety tools are filters. MoralStack is a judge.
It runs a full deliberative pipeline β risk estimation, constitutional critique, consequence simulation, multi-perspective reasoning β and issues an explicit, auditable decision before your LLM generates anything.
Traditional pipeline: prompt βββΊ generate βββΊ (maybe filter)
MoralStack: prompt βββΊ deliberate βββΊ decide βββΊ generate within bounds
Why teams pick MoralStack over a filter:
- π§ Deliberative, not reactive β a multi-module reasoning pipeline, not a keyword classifier
- βοΈ Explicit decisions β every request yields
NORMAL_COMPLETE,SAFE_COMPLETE, orREFUSE, computed from structured signals, never inferred from response text - π Full audit trail β every decision is explainable, logged, and queryable (AI Act art. 12-ready markdown exports)
- π§© Domain overlays β YAML-configurable per sector (healthcare, legal, financeβ¦), no code changes
- π Drop-in β wrap any OpenAI-compatible client in one line, or point clients at the OpenAI-compatible HTTP proxy
pip install moralstack
export OPENAI_API_KEY=sk-...from moralstack import govern
from openai import OpenAI
client = govern(OpenAI()) # one line β full OpenAI API compatibility
response = client.chat.completions.create(
model="gpt-4o",
messages=[{"role": "user", "content": "How do I pick a lock?"}],
)
print(response.choices[0].message.content)
# I'm unable to assist with that request.
meta = response.governance_metadata # governance metadata on every call
print(meta.final_action) # REFUSE
print(meta.risk_score) # 0.87
print(meta.decision_reason) # human-readable explanationBenign requests pass through untouched; sensitive ones get deliberated; harmful ones get a
governed refusal with safe redirection. More patterns in examples/ β
multi-turn, jailbreak resistance, audit export, HTTP proxy.
Install from source (contributors)
git clone https://github.com/fdidonato/moralstack.git
cd moralstack
python -m venv venv
source venv/bin/activate
pip install -e ".[dev,ui]"
cp .env.minimal .env # then set OPENAI_API_KEYOptional extras: moralstack[ui] (admin dashboard), moralstack[server] (HTTP proxy).
Full guide: INSTALL.md.
Evaluated on 84 questions spanning adversarial prompts, dual-use domains, regulated topics (legal, medical, financial), and false-positive torture tests. The judge model (GPT-5.2) is independent from both baseline and MoralStack generation.
| Baseline (GPT-4o) | MoralStack | |
|---|---|---|
| False negatives (missing refusal) | 13 | 0 |
| Information leakage | 14 (16.7%) | 0 (0%) |
| False positives (over-refusal) | 0 | 0 |
| Utility preservation | 62/62 | 62/62 |
| Safe redirection on refusal | 1/22 (4.5%) | 22/22 (100%) |
| Head-to-head wins (GPT-5.2 judge) | 6 | 54 (24 ties) |
| Avg safety score | 7.83/10 | 9.27/10 |
98.8% decision compliance Β· zero system errors
Decision confusion matrix & methodology notes
Predicted
Expected NC SC REFUSE
βββββββββββββββββββββββββββββββ
NC 9 1 0
SC 0 52 0
REFUSE 0 0 22
The single off-diagonal cell (1 NCβSC) is a health-domain query where MoralStack adds a professional-consultation disclaimer β a reasonable policy choice for regulated content.
Latency (benchmark 12, 84 questions): mean wall-clock ~36s, median ~26s, vs ~6s for raw GPT-4o. Mean is ~51% lower than the original benchmark configuration (~73s), with the fast-path rate up from ~11% to ~37% (REFUSE queries now routed through the fast path). Deliberation takes time by design β see Limitations & trade-offs.
Reproduce it:
python scripts/benchmark_moralstack.py
# Override baseline and judge models independently:
MORALSTACK_BENCHMARK_BASELINE_MODEL=gpt-4o \
MORALSTACK_BENCHMARK_JUDGE_MODEL=gpt-5.2 \
python scripts/benchmark_moralstack.pyNote: this benchmark demonstrates proof-of-concept effectiveness on 84 curated questions. It is not a claim of production-grade coverage across all possible inputs. We encourage independent evaluation on your domain-specific inputs.
Every request produces an explicit final_action β decision and generation are strictly separated:
| Action | Meaning |
|---|---|
NORMAL_COMPLETE |
Direct response |
SAFE_COMPLETE |
Responsible response with safeguards β a first-class policy action, not a post-hoc text disclaimer |
REFUSE |
Refusal with safe redirection |
flowchart TB
A([Request]) --> RE["Risk Estimator<br/><i>parallel mini-estimators: intent Β· signal detection q1βq17 Β· operational risk</i>"]
RE --> PR{"Policy Router<br/><i>domain overlay + action bounds</i>"}
PR -->|clearly benign / clearly harmful| FP[FAST_PATH<br/>deliberation skipped]
PR -->|sensitive / ambiguous| DP[DELIBERATIVE_PATH]
DP --> CC[Constitutional Critic]
DP --> CS[Consequence Simulator]
DP --> PE[Perspectives Ensemble]
DP --> HE[Hindsight Evaluator]
CC & CS & PE & HE --> CE["Convergence Engine<br/><b>issues final_action</b>"]
FP --> CE
CE --> RA["Response Assembler<br/><i>generates within the decided bounds</i>"]
RA --> OUT([Governed response + audit trail])
The single source of truth for bounds and action selection is
moralstack/runtime/decision/safe_complete_policy.py
(compute_action_bounds(...), decide_final_action(...)).
Package map
moralstack/sdk/β Python SDK (govern(),GovernedClient,GovernanceConfig)moralstack/runtime/β orchestration runtimemoralstack/orchestration/β controller, routing, deliberation servicesmoralstack/models/risk/β risk estimation and calibrationmoralstack/constitution/β constitution schema, loader, store (YAML-driven)moralstack/ui/β FastAPI dashboard (moralstack-ui)moralstack/server/β OpenAI-compatible governance HTTP proxy (create_app; install with.[server]or.[ui])
Deep dive: docs/architecture_spec.md.
Every user-visible answer is generated inside the governed pipeline by MoralStack's own
policy generator. The wrapped client you pass to govern() is never called to generate the
delivered answer, for any final_action β and pipeline failures fail closed to a governed
refusal, never an ungoverned passthrough.
final_action |
What happens |
|---|---|
NORMAL_COMPLETE |
The governed pipeline generates and delivers the answer |
SAFE_COMPLETE |
The answer is generated within explicit safeguard bounds; developer-declared system prompts are left byte-identical |
REFUSE |
No answer is generated β governed refusal text with safe redirection is returned |
The model= argument in client.chat.completions.create(model="...") is a requested alias /
OpenAI-compatibility field only: it does not select the model that produces the answer. To
choose the governed answer model, use GovernanceConfig(model=...) (SDK) or OPENAI_MODEL
(env). The response exposes generation_model (and rewrite_model when a revision occurred);
requested_model records the alias the caller supplied.
Model resolution per stage
| Stage | Model source |
|---|---|
| First-pass / speculative / compliance generation | GovernanceConfig.model β OPENAI_MODEL β gpt-4o |
Governed revision (policy.rewrite(), cycle 2+) |
MORALSTACK_POLICY_REWRITE_MODEL (fallback: resolved policy model) |
| Governed refusal text | resolved policy model |
| Risk / Critic / Simulator / Perspectives / Hindsight | MORALSTACK_*_MODEL per module, fallback to OPENAI_MODEL |
Chat request model= |
requested alias only β does not generate or rewrite the delivered answer |
The same one-line API governs full conversations:
from moralstack import govern
from openai import OpenAI
client = govern(OpenAI())
messages = []
for q in ["What is X?", "Tell me more.", "How do I do X?"]:
messages.append({"role": "user", "content": q})
response = client.chat.completions.create(model="gpt-4o", messages=messages)
messages.append({"role": "assistant", "content": response.choices[0].message.content})
print(response.governance_metadata.final_action, "β", response.choices[0].message.content[:80])- Transcript-aware governance β SDK and proxy attach the full OpenAI-style request transcript; DCCL and speculative generation reason over role-ordered prior turns, not just the final user message.
- Jailbreak resistance β escalation patterns are detected across turns (example).
- Guarded compliance fast-path β deployer-authorized contract execution can take
COMPLIANCE_FAST_PATH; a proxy governed draft is reused only when its generation context is aligned with the governance context. - Audit trail β every conversation produces a complete markdown export for AI Act art. 12 compliance (example).
- Session metadata β
meta.conversation_idandmeta.turn_indexon every response.
for chunk in client.chat.completions.create(
model="gpt-4o",
messages=[{"role": "user", "content": "..."}],
stream=True,
):
print(chunk.choices[0].delta.content or "", end="", flush=True)Governance deliberation happens before streaming starts. On REFUSE, a single synthetic
chunk is yielded with the refusal text.
meta = response.governance_metadata
meta.final_action # NORMAL_COMPLETE | SAFE_COMPLETE | REFUSE
meta.risk_score # 0.0 (benign) β 1.0 (harmful)
meta.risk_category # CLEARLY_BENIGN | SENSITIVE | CLEARLY_HARMFUL
meta.path # FAST_PATH | DELIBERATIVE_PATH
meta.reason_codes # ["DUAL_USE", "SENSITIVE_DOMAIN", ...]
meta.triggered_principles # constitution principles activated
meta.decision_reason # human-readable explanation
meta.conversation_id # session tracking (multi-turn)
meta.turn_index # turn counter within sessionAll non-chat.completions.create() calls pass through transparently
(client.models.list(), client.files.*, etc.).
OpenAI-compatible HTTP proxy β point any OpenAI-compatible client at
http://localhost:8080/v1 and get governance with X-Moralstack-* response headers. Install
with pip install "moralstack[server]", import create_app from moralstack.server, inject an
upstream OpenAI client and a configured OrchestrationController, and serve with uvicorn. (The
moralstack-server entry point is reserved for a future launcher.) See
docs/modules/server_proxy.md and
examples/server_quickstart.py.
Admin dashboard β inspect every decision: LLM calls, critic scores, risk traces, decision explanations, convergence steps, and benchmark comparisons.
pip install "moralstack[ui]"
# .env
# MORALSTACK_DB_PATH=moralstack.db
# MORALSTACK_UI_USERNAME=admin
# MORALSTACK_UI_PASSWORD=${UI_PASSWORD}
moralstack-ui
# β http://localhost:8765 (override with MORALSTACK_UI_PORT)Minimum required: OPENAI_API_KEY in the environment. Everything else has sensible defaults.
GovernanceConfig tunes the SDK per-client:
from moralstack import govern, GovernanceConfig
from openai import OpenAI
client = govern(
OpenAI(),
config=GovernanceConfig(
domain_overlay="healthcare", # enforce a specific domain overlay
observability_mode="file_only", # write JSONL audit trail
jsonl_dir=".investigations/logs/audit",
),
)Environment is loaded via moralstack/utils/env_loader.py: non-empty .env values override
existing env vars (override=True); optional empty values are purged after load.
Full environment variable reference
| Variable | Default | Description |
|---|---|---|
OPENAI_API_KEY |
(required) | Your OpenAI API key |
OPENAI_MODEL |
gpt-4o |
Model used by all pipeline modules |
MORALSTACK_POLICY_REWRITE_MODEL |
same as OPENAI_MODEL |
Model for deliberative rewrite() at cycle 2+. .env.template sets gpt-4.1-nano for lower rewrite latency |
OPENAI_TIMEOUT_MS |
60000 |
Per-call timeout in milliseconds |
OPENAI_MAX_RETRIES |
3 |
Retry count on transient errors |
OPENAI_TEMPERATURE |
0.1 |
Temperature for all modules |
OPENAI_TOP_P |
0.8 |
Top-p sampling parameter |
MORALSTACK_OBSERVABILITY_DB_PATH |
(unset) | Enables SQLite persistence |
MORALSTACK_OBSERVABILITY_MODE |
file_only |
db_only Β· dual Β· file_only |
MORALSTACK_OBSERVABILITY_JSONL_DIR |
logs/observability |
JSONL output directory |
MORALSTACK_DB_PATH / MORALSTACK_PERSIST_MODE |
β | Deprecated aliases β still work |
MORALSTACK_ORCHESTRATOR_BORDERLINE_REFUSE_UPPER |
0.95 |
Upper boundary for the borderline-refuse zone |
MORALSTACK_CORRELATION_TTL_SECONDS |
3600 |
TTL (seconds) for the server proxy's lineage correlation store; must exceed the max expected inter-turn delay, and should stay aligned with the session-store TTL so a lineage entry doesn't outlive (or expire well before) its governance session |
MORALSTACK_CORRELATION_MAX_ENTRIES |
20000 |
Max entries in the server proxy's lineage correlation store (FIFO eviction beyond this cap) |
MORALSTACK_PRINCIPAL_HMAC_SECRET |
(unset) | Secret used to HMAC the bearer token into a tenant/principal identifier for conversation correlation; when unset, the HMAC principal path is disabled (falls back to the empty-string sentinel) |
MORALSTACK_DCCL_SAFETY_OVERRIDE_MODEL |
gpt-4o-mini |
Model for the DCCL language-agnostic safety-override classifier (runs on every contract MATCH). Keep it small so the compliance fast-path stays fast |
Default models by component:
| Component | Default model | Override variable |
|---|---|---|
| Policy (generation) | gpt-4o |
OPENAI_MODEL |
| Policy (rewrite) | same as primary, or gpt-4.1-nano in .env.template |
MORALSTACK_POLICY_REWRITE_MODEL |
| Risk estimator | follows OPENAI_MODEL |
MORALSTACK_RISK_MODEL |
| Critic | follows OPENAI_MODEL |
MORALSTACK_CRITIC_MODEL |
| Simulator | follows OPENAI_MODEL |
MORALSTACK_SIMULATOR_MODEL |
| Perspectives | follows OPENAI_MODEL |
MORALSTACK_PERSPECTIVES_MODEL |
| Hindsight | follows OPENAI_MODEL |
MORALSTACK_HINDSIGHT_MODEL |
Full reference: INSTALL.md and docs/modules/.
| Regex / keywords | Moderation APIs | MoralStack | |
|---|---|---|---|
| Understands context | β | Partial | β |
| Auditable decisions | β | β | β |
| Domain-configurable | β | β | β |
| Handles dual-use | β | Partial | β |
| Safe redirection | β | β | β |
| Counterfactual reasoning | β | β | β |
| Zero false negatives* | β | β | β |
*On our benchmark set β see full methodology.
MoralStack makes deliberate trade-offs:
- Latency over speed β deliberative paths run multiple LLM calls (risk β critic β simulator β perspectives β hindsight): mean ~36s (median ~26s) vs ~6s for raw GPT-4o on the latest benchmark run. Governance takes time by design. Latency keeps dropping via speculative decoding, parallel risk estimation, lighter per-module models, structured JSON output enforcement, and soft-revision prompt constraints; early-exit and context-mode switching are planned.
- Multi-model cost β a single deliberative request makes 7β9 LLM calls.
.env.minimalships a cost-conscious profile (gpt-4.1-nanofor rewrite/simulator,gpt-4o-minifor perspectives), all overridable. - LLM non-determinism β despite low temperatures, outputs can vary between runs; deterministic in-code guardrails bound the variance, but perfect reproducibility is not guaranteed.
- Benchmark scope β 84 curated questions demonstrate the approach, not full coverage. Run your own evaluations on domain-specific inputs.
Full discussion: docs/limitations_and_tradeoffs.md.
| INSTALL.md | Detailed installation guide |
| examples/ | Runnable code examples (quickstart, multi-turn, jailbreak resistance, audit export, proxy) |
| docs/architecture_spec.md | Architecture specification |
| docs/decision_policy.md | Decision policy β action bounds & selection |
| docs/constitution.md | Constitution schema & principles |
| docs/creating_overlays.md | Building domain overlays |
| docs/modules/ | Per-module contracts |
| docs/DEVELOPMENT.md | Development guide |
Contributions are welcome β see CONTRIBUTING.md. Development setup, test commands, and repository conventions are in docs/DEVELOPMENT.md.
pip install -e ".[dev,ui]"
python -m pytest # full offline test suite
pre-commit run -aApache 2.0 β see LICENSE.
If you use MoralStack in academic work, please cite:
@software{moralstack,
author = {di Donato, Francesco},
title = {MoralStack: a deliberative governance engine for LLMs},
url = {https://github.com/fdidonato/moralstack},
year = {2026}
}