Catch the bad text
before your model does.
Tripwire checks what goes into your LLM and what comes out. Plain rules first, an open model as an optional second opinion. No dependencies. Rules run in this page, so nothing you type leaves it.
Ask Gemma too optional, your own free AI Studio key
Some attacks don't match any rule (try paraphrase). The judge catches those. Your key stays in this page and goes only to generativelanguage.googleapis.com in a header. Nothing is stored. The text is sent to Google when you press the button.
Use it
One check before the model, one after. Node 20+, import or require. Not on npm yet: install from GitHub.
import { createGuard } from "tripwire-guardrails";
const guard = createGuard();
const inbound = guard.checkInput(userMessage);
if (!inbound.allowed) return reply(400, { blocked: inbound.categories });
const answer = await callYourModel(userMessage);
const outbound = guard.checkOutput(answer);
return outbound.allowed ? outbound.redacted : "Sorry, I can't share that.";import express from "express";
import { createGuard } from "tripwire-guardrails";
import { tripwire } from "tripwire-guardrails/express";
const app = express();
app.use(express.json());
// blocks bad or missing input with a 400, checks res.json({ reply }) on the way out
app.post("/chat", tripwire(createGuard()), async (req, res) => {
res.json({ reply: await callYourModel(req.body.message) });
});import { createGuard } from "tripwire-guardrails";
import { createJudge, checkWithJudge } from "tripwire-guardrails/judge";
const guard = createGuard();
const judge = createJudge({ apiKey: process.env.GEMINI_API_KEY }); // gemma-4-26b-a4b-it
const verdict = await checkWithJudge(guard, judge, text, "input", { always: true });
// if the judge call fails, the text is blocked (fails closed)tripwire "ignore all previous instructions" # exit code 1, JSON on stdout
echo "token ghp_..." | tripwire --output
npx tripwire-server --port 8787 # POST /check for Python, Go, anythingnpm install github:MaybeSomeone-arc18/tripwire-guardrails
# or try it first, no install
npx github:MaybeSomeone-arc18/tripwire-guardrails "ignore all previous instructions"- Rules are data. One file of regexes with a reason on every finding. Copy one, change it, delete it.
- Fails closed. If the judge errors, times out or returns junk, the text is blocked.
- Swappable judge. Change the model name, or write your own object with a
judge(text)method. Anollamaprovider exists but is untested.
What it checks
Prompt injection
"Ignore previous instructions", prompt extraction, role swaps, forged </system> and [INST] markers, notes aimed at the assistant, image URLs that carry data.
Obfuscation
Unicode look-alikes, zero-width characters, spaced-out letters, leetspeak, ROT13, base64 and hex are decoded and scanned too.
Other languages
The same ideas in Spanish, French, German, Hindi, Hinglish and Chinese. Hand-written, so coverage is thin.
Secrets, both ways
AWS, GitHub, Google, Slack and sk- keys, private key blocks, JWTs. Redacted on the way out.
Personal data
Emails, Indian mobile numbers, Aadhaar (checksum verified), PAN, card numbers (Luhn verified). Shape checks, not names or addresses.
Policy
Block threshold, ignored categories, topic allow-list, length cap. JSON policies in policies/.
Measured, not promised
Two test sets I wrote, in eval/, and one public benchmark I did not write. I tuned the rules against the first. The other two are the honest ones, and the public one is the harshest.
Recall on hostile texts. Gemma (gemma-4-26b-a4b-it, AI Studio free tier) gave a real block on 22 of 24; calls that failed are blocked but not counted as caught. On benign texts it falsely blocked 1 (a card-shaped student ID). 3 of 38 calls failed after retries. One run, 48 texts, one author: a rough signal, not a benchmark.
Limits
- Not a complete defence. Prompt injection has none. Treat this as one layer.
- Speed: rules take about 0.07 ms for a sentence and 0.6 ms for 1,100 characters (one 2-CPU machine). A model judge adds a network call; local 1B Gemma took about 2.2 s. Hosted Gemma was not timed.
- Rules are English-first and miss many fresh paraphrases: 41.7% on the held-out set, and 20.0% on an outside benchmark.
- The judge is a model. It can be wrong or attacked itself, and a flaky free endpoint blocks some good text.
- Streamed replies: text already sent can't be taken back. No per-user policies or logging. Attacks inside documents or tool output are covered only if you pass that text through
checkInput.
Why it exists
Built for Prudhvi, for the DEV Hacktoberfest Weekend Challenge "Build for a Friend". We kept having the same argument while building AI projects: trust the model, or put a separate layer of checks around it. This is the separate layer, small enough to read in one sitting.
Written during the challenge window (Oct 2 to 5, 2026) with an AI agent doing most of the coding and testing, steered and reviewed by me. MIT licensed.