tripwire

Catch the bad text
before your model does.

Tripwire checks what goes into your LLM and what comes out. Plain rules first, an open model as an optional second opinion. No dependencies. Rules run in this page, so nothing you type leaves it.

Direction
Ask Gemma too optional, your own free AI Studio key

Some attacks don't match any rule (try paraphrase). The judge catches those. Your key stays in this page and goes only to generativelanguage.googleapis.com in a header. Nothing is stored. The text is sent to Google when you press the button.

Use it

One check before the model, one after. Node 20+, import or require. Not on npm yet: install from GitHub.

import { createGuard } from "tripwire-guardrails";
const guard = createGuard();

const inbound = guard.checkInput(userMessage);
if (!inbound.allowed) return reply(400, { blocked: inbound.categories });

const answer = await callYourModel(userMessage);
const outbound = guard.checkOutput(answer);
return outbound.allowed ? outbound.redacted : "Sorry, I can't share that.";

What it checks

Prompt injection

"Ignore previous instructions", prompt extraction, role swaps, forged </system> and [INST] markers, notes aimed at the assistant, image URLs that carry data.

Obfuscation

Unicode look-alikes, zero-width characters, spaced-out letters, leetspeak, ROT13, base64 and hex are decoded and scanned too.

Other languages

The same ideas in Spanish, French, German, Hindi, Hinglish and Chinese. Hand-written, so coverage is thin.

Secrets, both ways

AWS, GitHub, Google, Slack and sk- keys, private key blocks, JWTs. Redacted on the way out.

Personal data

Emails, Indian mobile numbers, Aadhaar (checksum verified), PAN, card numbers (Luhn verified). Shape checks, not names or addresses.

Policy

Block threshold, ignored categories, topic allow-list, length cap. JSON policies in policies/.

Measured, not promised

Two test sets I wrote, in eval/, and one public benchmark I did not write. I tuned the rules against the first. The other two are the honest ones, and the public one is the harshest.

Tuned set, rules alone (39 hostile, 50 benign, 0 false positives)100%
Held-out set, rules alone (24 hostile, 24 benign, 0 false positives)41.7%
Outside benchmark (deepset test split), rules alone (60 hostile, 56 benign, 0 false positives; 1.7% before my fixes)20.0%
Held-out set, rules plus Gemma91.7%

Recall on hostile texts. Gemma (gemma-4-26b-a4b-it, AI Studio free tier) gave a real block on 22 of 24; calls that failed are blocked but not counted as caught. On benign texts it falsely blocked 1 (a card-shaped student ID). 3 of 38 calls failed after retries. One run, 48 texts, one author: a rough signal, not a benchmark.

Limits

Why it exists

Built for Prudhvi, for the DEV Hacktoberfest Weekend Challenge "Build for a Friend". We kept having the same argument while building AI projects: trust the model, or put a separate layer of checks around it. This is the separate layer, small enough to read in one sitting.

Written during the challenge window (Oct 2 to 5, 2026) with an AI agent doing most of the coding and testing, steered and reviewed by me. MIT licensed.