Pilot · ShieldGemma 2B · 30 harmful + 30 safe prompts

Type a harmful prompt in Hinglish and 1 in 5 gets past the AI guardrail.

Kavach rewrites code-mixed text into plain English before the guardrail reads it. Same guardrail, no retraining. In our pilot the slip rate fell from 20% to 0%.

One click picks a sample sentence and runs the shield on it. Nothing is ever answered, only rewritten.

20%Heavy Hinglish, no shieldof harmful prompts slip through
0%Heavy Hinglish, with Kavachof harmful prompts slip through
+0.6 sThe costper prompt, and a few more safe prompts wrongly blocked

How it works

Kavach is a small shield that sits in front of any guardrail. It never answers a prompt. It only rewrites it.

1

A prompt arrives in Hinglish

Hindi words in Latin script with loose spelling. A guardrail trained mostly on English reads it poorly.

2

The shield rewrites it in English

One instruction to a small model: rewrite this in plain English, keep the meaning, do not answer it.

3

The guardrail scores both

The higher score wins. If the shield cannot rewrite the text, the prompt is blocked. It fails closed.

Try the shield

Pick a sample, or type any Hinglish sentence. You will see what the guardrail read before and what it reads after Kavach.

Risky prompts

Harmful requests. Click one and watch the shield expose the intent in plain English, so the guardrail can block it.

Safe prompts

Harmless requests. Click one and watch it pass through the shield, and see where the guardrail still over-blocks.

Pilot results

Same prompts, four conditions. Heavy Hinglish is where the guardrail fails, and where Kavach brings it back.

See all 60 prompts and their verdicts

Harmful prompt text is never shown here, only its category and the guardrail's verdict. Safe prompts are shown in full.