Aiso data practice

Inside Rampart: the 14.7 MB privacy model behind safer AI conversations

Benjamin Tannenbaum, Founder and CEO, Aiso
By Benjamin Tannenbaum · Founder and CEO, Aiso · LinkedIn
6 min read

First published .

America.gov's new AI makes one privacy behavior obvious: try to submit personal information and it is filtered before the conversation continues. Behind that approach is Rampart, a tiny local model built by the U.S. National Design Studio. We use the same technology as one of the tools for cleaning Aiso's dataset of real AI-human conversations.

Rampart illustration from National Design Studio
Rampart illustration. Source: National Design Studio.

America.gov launched on September 29 as a new AI front door to U.S. government information and services. One of the more interesting product details is not the chatbot itself. It is what happens before a message reaches it.

Enter personal information and the interface can filter it before the message is sent onward. The technology behind that privacy pattern is Rampart, an open-source system built by the U.S. National Design Studio for removing personally identifiable information locally.

A post from Pliny the Liberator highlighted the implementation. It caught my attention because this is not just an interesting government AI feature. We already use Rampart as one of the tools for cleaning Aiso's dataset of real AI-human conversations.

The clever part is where it runs. Rampart was designed so the privacy step can happen on the user's device, before raw text reaches a remote model, server or logging system. The shipped model is only 14.7 MB.

14.7 MB
Shipped model size, including tokenizer
98.42%
Private-term recall reported on a 30,000-row held-out test
3.9 ms
Reported p50 browser latency with WebGPU

What Rampart actually does, in simple terms

Rampart uses two readers. The first is conventional software: rules and validation for information with recognizable structure, such as email addresses, Social Security numbers, phone numbers, cards, routing details and IP addresses.

The second is a small MiniLM language model. It handles the messier cases where context matters, especially names and street addresses. Detected values are replaced with placeholders before the text continues through the rest of the pipeline. The remote AI can work with the meaning of the request without receiving the original personal value.

National Design Studio reports 98.42% private-term recall across seven supported Latin-script languages on a 30,000-row held-out OpenPII test set. The same published benchmark reports 3.9 ms median browser latency using WebGPU.

Those numbers are useful, but they are not a privacy guarantee. The Rampart model card explicitly describes it as a first line of defense, and its own evaluation still contains missed private terms. That is the right way to think about it: a very useful layer, not a reason to stop thinking about data minimization and downstream controls.

We use the same technology when cleaning Aiso's conversation data

Aiso works with a dataset of real AI-human conversations. The useful part for marketers is the language: what people ask first, what they ask next, which objections appear, how they compare options and when a brand enters the conversation.

A person's name, email address, phone number or street address does not make that analysis better. It is noise for the job we are trying to do.

So Rampart is one of the tools we use in the cleaning layer for our conversation dataset. At scale, the goal is to remove personal information while preserving the parts that matter for analysis: the question, the context, the follow-up and the sequence of the conversation.

That is an important distinction. "Real conversation data" should not mean "keep every field because it might be useful later." If the research question is about demand, language and recommendation behavior, identity is usually unnecessary.

What this means for marketers and SEO teams

AI search gives marketers a much richer signal than a keyword. A person can start broad, add a constraint, reject an option, mention a budget and only then ask for a recommendation. If you only study the opening prompt, you miss a lot of the demand.

That richer data is valuable because it shows how a decision develops. It also makes privacy engineering more important. People tell conversational systems things they would never type into a traditional search box. The useful marketing signal is the intent and the progression of the conversation, not the identity of the person having it.

For SEO and AI-search research, that leads to a useful rule: preserve the language and intent you need, remove identity you do not need, and do the removal as early in the pipeline as possible.

Small local models are becoming infrastructure

Rampart is also a useful example of where small models can beat the instinct to send every task to a frontier LLM. PII detection is narrow. It needs low latency, predictable behavior and the ability to run close to the data.

A 14.7 MB model that can run in a browser changes the architecture. The privacy step does not need to become another remote API call that sees the raw text first.

For product teams, researchers and marketers working with AI conversation data, that is the broader lesson. Some of the most useful AI infrastructure will be small, specialized and almost invisible.

Sources