arlong
← All field notesSecurity / Research preview

INTRODUCING
ARLONG
BOUNDARY 1.

A stricter boundary between
untrusted sources and agent context.

The web can contain useful evidence and hostile instructions on the same page. Boundary 1 adds source screening, quarantine decisions and delivery checks to Arlong’s API and MCP tools.

Source text is evidence.
It cannot grant authority.

Instructions embedded in a page or a tool response can try to redirect an agent’s task, impersonate a trusted role, or request access to private information. Boundary 1 evaluates external content before it enters Arlong’s research evidence.

01

Screen before synthesis

Search previews and downloaded content are checked before candidate scouting and evidence synthesis. Retained text is scanned in overlapping windows to cover instructions that appear late in a document.

02

Remove hostile sections

During extraction, active and hidden markup is removed, suspect text sections are isolated, and the retained content is screened again as a whole. Useful evidence can survive without carrying the detected instructions. Content that cannot be safely isolated stays quarantined. The screening tool separately reports a context decision for the submitted payload.

03

Check the delivery boundary

API and hosted MCP evidence responses receive a final delivery check. Rejected text is withheld from both the structured result and the MCP text representation. Delivery-check failures return an error rather than the unchecked result.

Boundary 1 is a server-side security system, not a newly trained language model. Its implementation remains private. We publish benchmark provenance and measured results while keeping the internal policy implementation private.

Measure the gap.
Then reduce it.

We copied a pinned, MIT-licensed snapshot of InjecAgent, a research benchmark for indirect prompt injection in tool-integrated agents. We screened the complete tool-response text in all four base and enhanced dataset files. No attacker tools were executed.

2108Attack payloads tested
95.83%Payloads blocked
+45.83Percentage points vs previous
◆ Arlong Boundary 1● Previous scanner

Attack payload block rate · higher is better

Boundary 1 and previous scanner by InjecAgent settingdh_base: Boundary 1 82.75%, previous 0.0%. dh_enhanced: Boundary 1 100.0%, previous 100.0%. ds_base: Boundary 1 100.0%, previous 0.0%. ds_enhanced: Boundary 1 100.0%, previous 100.0%. Settings are categorical, not cost or a progression of difficulty. 0%20%40%60%80%100% 82.75%HarmBase · n=510100.0%HarmEnhanced · n=510100.0%Data theftBase · n=544100.0%Data theftEnhanced · n=544InjecAgent dataset setting

Four categorical settings. Lines connect measured points for comparison; they do not represent cost, difficulty or additional experiments.

Results by dataset setting
SettingCasesPrevious blockedBoundary 1 blocked
DH BASE5100422
DH ENHANCED510510510
DS BASE5440544
DS ENHANCED544544544

Boundary 1 blocked 2020 of 2108 payloads, compared with 1054 for the previous scanner. It blocked none of 17 deduplicated, neutralized tool-response templates. Those controls are small and synthetic; they do not establish a production false-positive rate.

Median local screening time was 0.69 ms per case; the 95th percentile was 1.16 ms. These measurements exclude network time, model calls and the rest of a research run.

Evaluation methodology and supporting checks

Upstream revision: f19c9f2c79a41046eb13c03c51a24c567a8ffa07. Baseline scanner: f61ad41. This offline classification evaluation covers complete tool-response text. The public corpus informed policy development; it is a regression evaluation, not a held-out test. Base and enhanced variants share underlying scenarios.

88 attack payloads passed screening. This score measures classification, not live-agent attack success or protection against adaptive, multilingual or multimodal attacks. The 17 neutralized controls are synthetic and do not establish a production false-positive rate.

A separate post-development check used the PINT repository’s eight public examples, without further policy tuning: 1/2 attack examples blocked and 0/6 benign examples quarantined. This is an example check, not a full PINT score.

Passing content remains untrusted. Boundary 1 screens Arlong evidence; agent applications retain responsibility for scoped permissions and action authorization. The research preview does not guarantee immunity to prompt injection. Oversized inputs are rejected rather than silently approved.

Run python -m benchmarks.boundary1 against the pinned corpus. The report records hashes, counts and methodology. Timing excludes network and model calls.

Download evaluation report ↓

Research beyond
the client timeout.

Deep research can now continue on the server after an authenticated request returns. Set cloud_run: true to receive a request ID and estimated completion time immediately. Use request_id_search for progress and the full answer when the job finishes.

Polling costs no credits. Results are account-scoped and retained for seven days. Boundary delivery checks apply when cloud evidence is retrieved.

Start at the tool boundary.

Use arlong_screen before adding external text to an agent’s context. Check boundary.allowed_for_context; ingest only the returned safe_content. An allow decision is not authorization to execute instructions found in that text.

{
  "name": "arlong_screen",
  "arguments": {
    "text": "External text to evaluate",
    "raw_text": "Original markup, when available"
  }
}