Source text is evidence.
It cannot grant authority.
Instructions embedded in a page or a tool response can try to redirect an agent’s task, impersonate a trusted role, or request access to private information. Boundary 1 evaluates external content before it enters Arlong’s research evidence.
Screen before synthesis
Search previews and downloaded content are checked before candidate scouting and evidence synthesis. Retained text is scanned in overlapping windows to cover instructions that appear late in a document.
Remove hostile sections
During extraction, active and hidden markup is removed, suspect text sections are isolated, and the retained content is screened again as a whole. Useful evidence can survive without carrying the detected instructions. Content that cannot be safely isolated stays quarantined. The screening tool separately reports a context decision for the submitted payload.
Check the delivery boundary
API and hosted MCP evidence responses receive a final delivery check. Rejected text is withheld from both the structured result and the MCP text representation. Delivery-check failures return an error rather than the unchecked result.
Boundary 1 is a server-side security system, not a newly trained language model. Its implementation remains private. We publish benchmark provenance and measured results while keeping the internal policy implementation private.
Measure the gap.
Then reduce it.
We copied a pinned, MIT-licensed snapshot of InjecAgent, a research benchmark for indirect prompt injection in tool-integrated agents. We screened the complete tool-response text in all four base and enhanced dataset files. No attacker tools were executed.
Attack payload block rate · higher is better
Four categorical settings. Lines connect measured points for comparison; they do not represent cost, difficulty or additional experiments.
| Setting | Cases | Previous blocked | Boundary 1 blocked |
|---|---|---|---|
| DH BASE | 510 | 0 | 422 |
| DH ENHANCED | 510 | 510 | 510 |
| DS BASE | 544 | 0 | 544 |
| DS ENHANCED | 544 | 544 | 544 |
Boundary 1 blocked 2020 of 2108 payloads, compared with 1054 for the previous scanner. It blocked none of 17 deduplicated, neutralized tool-response templates. Those controls are small and synthetic; they do not establish a production false-positive rate.
Median local screening time was 0.69 ms per case; the 95th percentile was 1.16 ms. These measurements exclude network time, model calls and the rest of a research run.
Evaluation methodology and supporting checks
Upstream revision: f19c9f2c79a41046eb13c03c51a24c567a8ffa07. Baseline scanner: f61ad41. This offline classification evaluation covers complete tool-response text. The public corpus informed policy development; it is a regression evaluation, not a held-out test. Base and enhanced variants share underlying scenarios.
88 attack payloads passed screening. This score measures classification, not live-agent attack success or protection against adaptive, multilingual or multimodal attacks. The 17 neutralized controls are synthetic and do not establish a production false-positive rate.
A separate post-development check used the PINT repository’s eight public examples, without further policy tuning: 1/2 attack examples blocked and 0/6 benign examples quarantined. This is an example check, not a full PINT score.
Passing content remains untrusted. Boundary 1 screens Arlong evidence; agent applications retain responsibility for scoped permissions and action authorization. The research preview does not guarantee immunity to prompt injection. Oversized inputs are rejected rather than silently approved.
Run python -m benchmarks.boundary1 against the pinned corpus. The report records hashes, counts and methodology. Timing excludes network and model calls.
Research beyond
the client timeout.
Deep research can now continue on the server after an authenticated request returns. Set cloud_run: true to receive a request ID and estimated completion time immediately. Use request_id_search for progress and the full answer when the job finishes.
Polling costs no credits. Results are account-scoped and retained for seven days. Boundary delivery checks apply when cloud evidence is retrieved.
Start at the tool boundary.
Use arlong_screen before adding external text to an agent’s context. Check boundary.allowed_for_context; ingest only the returned safe_content. An allow decision is not authorization to execute instructions found in that text.
{
"name": "arlong_screen",
"arguments": {
"text": "External text to evaluate",
"raw_text": "Original markup, when available"
}
}