PII redaction before your data reaches any AI model

Worried about sending client data to an AI provider, or to a model hosted somewhere you don't know? Yellowjacket swaps personal details for placeholders before anything leaves our server, and puts the originals back in your answer.

What is hidden

  • Names of people and organisations
  • Email addresses and phone numbers
  • Street addresses
  • Bank accounts, BSB / sort / routing numbers, IBANs
  • Card numbers (Luhn-checked), tax and ID numbers (ABN, TFN, SSN, EIN, VAT)
  • Web addresses, IP addresses and secret keys

"Acme Pty Ltd" becomes {ORG_1} everywhere it appears, so the models can still match it across documents. Amounts, dates and invoice or PO numbers are kept, because the checks need them.

Does redaction hurt accuracy?

Barely. On 20 real contracts the same job scored 70/80 with redaction and 71/80 without.

More than redaction

  • Zero data retention: every model call is routed only to providers that keep nothing and don't train on it.
  • Read-only connectors: Drive, SharePoint, Gmail, Slack, Xero and databases are read, never written.
  • Your choice on training: nothing you send trains anything unless you switch it on.
  • A record per document: which model read it, which checks it passed.

No detector finds every name; redaction lowers exposure, it does not make it zero. Photos are sent as they are.

How it works

  1. A top model plans the job once. Claude Opus looks at a few samples and writes the instructions, the checks and the arithmetic. You pay for that once per kind of job, and the plan is kept for next time.
  2. Cheap models do the reading. A short tryout picks the cheapest open models that pass the job's checks; they read every document. Sums, dates and comparisons are done in code, not by a model.
  3. Every answer is checked. Totals have to reconcile and two different models have to agree. When they don't, a stronger model reads it again; only real disagreements go back to the top model.
  4. You get a receipt. Before the job you see what Claude, ChatGPT and Gemini would each charge alone; after it, what each document cost and which checks it passed.

Questions people ask

How do I redact PII before sending data to an LLM?

Detect names and identifiers (patterns for structured IDs, a named-entity model for names), replace each with a consistent placeholder, send the redacted text, and restore the placeholders in the reply. Yellowjacket does this on every job, chat and API call by default.

Is it safe to use Chinese or other open models with my data?

With redaction on, the models see placeholders instead of names and account numbers, and calls go only to providers with zero data retention.

Can I switch redaction off?

Yes, per workspace, in Settings, Privacy.

Take the sting out of your AI bill

Run a job