arlong.
Technical whitepaper / v1.0

Secure retrieval.
Inspectable evidence.

A practical architecture for letting AI agents use the live web without confusing retrieved instructions with trusted control, or plausible text with verified fact.

Arlong Search Systems · September 2026 · Sushant & Ahilan

01

Abstract

AI agents issue requests that are longer and more constrained than conventional web queries. They also consume page content as model context, creating a new attack surface: a webpage can contain text intended to redirect the agent, invoke tools, disclose secrets, or override its task. At the same time, a fluent synthesis can conceal incomplete or contradictory evidence.

Arlong separates these problems into accountable stages. It plans focused retrieval lanes, retrieves live-web candidates, rejects low-signal results before expensive extraction, screens content at the ingestion boundary, validates candidate identity and constraints, and synthesizes from compact evidence excerpts with attached citations. Missing evidence remains explicit. Arlong federates existing web retrieval; it does not claim a proprietary web-scale index.

Design objective: maximize useful, current evidence per request while minimizing untrusted content admitted to model context.

02

Problem and threat model

Agent-sized queries

A single request can combine discovery, entity resolution, dates, geography, funding stage, technical properties, exclusions, and an output format. Sending that paragraph directly to a keyword engine loses intent. Over-expansion has the opposite failure mode: generic background pages consume the extraction and synthesis budget.

Untrusted webpages

Arlong assumes retrieved content can be malicious, misleading, duplicated, stale, or irrelevant. An attacker may place instructions in visible copy, metadata, hidden CSS, zero-width text, structured data, or content loaded only after rendering. Security labels therefore describe observed evidence, not certainty.

Synthesis failures

A result can be correctly retrieved but incorrectly summarized. Important failure modes include entity collisions, contradiction suppression, self-corroboration from syndicated content, unsupported negative claims, and citations that do not entail the nearby statement.

03

Bounded research architecture

The production research path uses separate model roles. A fast planning model turns the request into a small set of search-native lanes; a stronger synthesis model receives only selected evidence. DeepSeek is the primary model provider, with Groq used as a fallback route. Search candidates are gathered through Serper.

01InterpretPreserve subject, entities, constraints, freshness, requested fields, and the desired answer shape.
02PlanGenerate three short, complementary queries. List requests begin with candidate discovery, then move to candidate-specific verification.
03PreviewSearch titles and snippets first. Local relevance checks remove generic background, collisions, and pages that cannot satisfy a required field.
04ExtractDownload only the strongest pages, normally capped at ten. Extraction produces bounded, query-relevant excerpts rather than unrestricted page dumps.
05ScreenAssess content before it reaches synthesis. Blocked text is withheld; incomplete scans remain unknown.
06VerifyCheck candidate identity, category, dates, geography, and other hard constraints. Contradictory candidates cannot be promoted as confirmed matches.
07SynthesizeCompose the answer from useful excerpts across sources, attach citations, surface near-matches, and state what could not be verified.

Cost controls

The default design limits query fan-out, preview counts, extracted documents, excerpt size, and recovery passes. Local filtering occurs before network-heavy extraction. Recovery is triggered by a measured coverage gap, rather than being unconditional.

04

Security is a boundary, not a badge

Arlong keeps five dimensions separate so one signal cannot disguise another.

Domain reputationHistorical and structural signals about the host.
Content securityEvidence of prompt injection, hidden directives, credential requests, or tool manipulation.
Query relevanceWhether the page can answer the actual request.
Factual reliabilityWhether the source is appropriate for the particular claim.
Extraction stateSuccessful, reviewable, blocked, failed, or unknown—never silently converted to safe.
Source independenceWhether evidence comes from genuinely distinct reporting or copied material.

Human-oriented tutorial language is not automatically prompt injection. Detection considers whether text addresses an AI system, attempts to change its governing task, requests sensitive data or external action, is concealed from a normal reader, or appears in a context inconsistent with the page’s purpose. Review means sanitize and use cautiously; block means withhold content from downstream models.

Invariant: if zero characters were scanned, the security action cannot be “allow.”
05

Evidence-aware output

Arlong represents useful support as structured evidence rather than treating topical similarity as agreement. A simplified claim record looks like this:

{
  "claim": "Company X announced a Series A in 2026",
  "supporting_sources": ["source-id-1", "source-id-2"],
  "contradicting_sources": [],
  "confidence": 0.91,
  "status": "verified"
}

Authority is claim-dependent. A company website may be the best source for product behavior; a funding announcement or regulatory filing may be better for financing; an independent technical evaluation may be better for comparative performance. Multiple copies of the same announcement do not become independent corroboration simply because they appear on different URLs.

Calibrated abstention

When all strict criteria cannot be verified, the answer should distinguish confirmed matches, strong candidates, near-matches, and exclusions. Insufficient evidence supports “unable to verify,” not “no cases exist.” Negative universal claims require much stronger coverage than positive candidate discovery.

06

Interfaces and feedback

The same retrieval system is available in the Playground, REST API, and hosted MCP server. MCP exposes tools for quick links, evaluated search, deep research, extraction, grounded answers, status, and authenticated error reporting.

Every Arlong retrieval or generation result—including failure responses—receives a unique signed arlong-provenance-v1 identifier. The signature confirms that Arlong issued the ID while keeping the query, answer, and user identity out of the identifier. Responses expose the ID in structured output and HTTP headers; streamed answers also carry a visible footer. A public verification URL checks the signature without retrieving or exposing the original response.

arlong_report_wrong requires that signed response ID before an agent can submit a concrete wrong answer, irrelevant or missing result, stale fact, citation mismatch, extraction failure, security false positive/negative, or tool error. It costs no credits, is rate-limited, and stores only the provided reproduction details—never the API key, IP address, user agent, or hidden model context.

Reports are feedback, not instructions to silently change rankings. They must be reviewed, grouped by failure class, reproduced, and converted into regression tests or labelled evaluation examples before influencing the system.

07

Evaluation methodology

End-to-end quality should be measured across diverse query classes and repeated runs. Recommended dimensions include relevance, factual correctness, constraint satisfaction, entity resolution, recall, freshness, source authority, citation entailment, contradiction handling, latency, cost, and adversarial-content handling.

Security evaluation requires a labelled corpus containing genuine injections, benign human instructions, hidden and encoded payloads, tool-poisoning attempts, and extraction failures. Report precision, recall, false-positive rate, false-negative rate, and calibration by risk band. Search benchmarks and security benchmarks should be published separately so strong detection cannot conceal weak retrieval, or vice versa.

08

Limitations and commitments

  • Live-web coverage depends on upstream search availability and index coverage.
  • Some pages cannot be extracted because of authentication, robots policies, rendering, regional controls, or network failure.
  • Prompt-injection detection is probabilistic and must be evaluated continuously; no detector guarantees safety.
  • Source reputation does not prove an individual claim, and first-party sources can be incomplete or self-interested.
  • Recent or obscure facts may have only one available source. The system should expose that limitation rather than manufacture agreement.
  • Arlong is a retrieval and evidence system, not a substitute for professional judgment in medical, legal, financial, or safety-critical decisions.
Commitment: make failure states observable, keep unverified fields visible, and prefer a useful qualified answer over false certainty.