Abstract
AI agents issue requests that are longer and more constrained than conventional web queries. They also consume page content as model context, creating a new attack surface: a webpage can contain text intended to redirect the agent, invoke tools, disclose secrets, or override its task. At the same time, a fluent synthesis can conceal incomplete or contradictory evidence.
Arlong separates these problems into accountable stages. It plans focused retrieval lanes, retrieves live-web candidates, rejects low-signal results before expensive extraction, screens content at the ingestion boundary, validates candidate identity and constraints, and synthesizes from compact evidence excerpts with attached citations. Missing evidence remains explicit. Arlong federates existing web retrieval; it does not claim a proprietary web-scale index.
Design objective: maximize useful, current evidence per request while minimizing untrusted content admitted to model context.
Problem and threat model
Agent-sized queries
A single request can combine discovery, entity resolution, dates, geography, funding stage, technical properties, exclusions, and an output format. Sending that paragraph directly to a keyword engine loses intent. Over-expansion has the opposite failure mode: generic background pages consume the extraction and synthesis budget.
Untrusted webpages
Arlong assumes retrieved content can be malicious, misleading, duplicated, stale, or irrelevant. An attacker may place instructions in visible copy, metadata, hidden CSS, zero-width text, structured data, or content loaded only after rendering. Security labels therefore describe observed evidence, not certainty.
Synthesis failures
A result can be correctly retrieved but incorrectly summarized. Important failure modes include entity collisions, contradiction suppression, self-corroboration from syndicated content, unsupported negative claims, and citations that do not entail the nearby statement.
Bounded research architecture
The production research path uses separate model roles. A fast planning model turns the request into a small set of search-native lanes; a stronger synthesis model receives only selected evidence. DeepSeek is the primary model provider, with Groq used as a fallback route. Search candidates are gathered through Serper.
Cost controls
The default design limits query fan-out, preview counts, extracted documents, excerpt size, and recovery passes. Local filtering occurs before network-heavy extraction. Recovery is triggered by a measured coverage gap, rather than being unconditional.
Security is a boundary, not a badge
Arlong keeps five dimensions separate so one signal cannot disguise another.
Human-oriented tutorial language is not automatically prompt injection. Detection considers whether text addresses an AI system, attempts to change its governing task, requests sensitive data or external action, is concealed from a normal reader, or appears in a context inconsistent with the page’s purpose. Review means sanitize and use cautiously; block means withhold content from downstream models.
Evidence-aware output
Arlong represents useful support as structured evidence rather than treating topical similarity as agreement. A simplified claim record looks like this:
{
"claim": "Company X announced a Series A in 2026",
"supporting_sources": ["source-id-1", "source-id-2"],
"contradicting_sources": [],
"confidence": 0.91,
"status": "verified"
}Authority is claim-dependent. A company website may be the best source for product behavior; a funding announcement or regulatory filing may be better for financing; an independent technical evaluation may be better for comparative performance. Multiple copies of the same announcement do not become independent corroboration simply because they appear on different URLs.
Calibrated abstention
When all strict criteria cannot be verified, the answer should distinguish confirmed matches, strong candidates, near-matches, and exclusions. Insufficient evidence supports “unable to verify,” not “no cases exist.” Negative universal claims require much stronger coverage than positive candidate discovery.
Interfaces and feedback
The same retrieval system is available in the Playground, REST API, and hosted MCP server. MCP exposes tools for quick links, evaluated search, deep research, extraction, grounded answers, status, and authenticated error reporting.
Every Arlong retrieval or generation result—including failure responses—receives a unique signed arlong-provenance-v1 identifier. The signature confirms that Arlong issued the ID while keeping the query, answer, and user identity out of the identifier. Responses expose the ID in structured output and HTTP headers; streamed answers also carry a visible footer. A public verification URL checks the signature without retrieving or exposing the original response.
arlong_report_wrong requires that signed response ID before an agent can submit a concrete wrong answer, irrelevant or missing result, stale fact, citation mismatch, extraction failure, security false positive/negative, or tool error. It costs no credits, is rate-limited, and stores only the provided reproduction details—never the API key, IP address, user agent, or hidden model context.
Reports are feedback, not instructions to silently change rankings. They must be reviewed, grouped by failure class, reproduced, and converted into regression tests or labelled evaluation examples before influencing the system.
Evaluation methodology
End-to-end quality should be measured across diverse query classes and repeated runs. Recommended dimensions include relevance, factual correctness, constraint satisfaction, entity resolution, recall, freshness, source authority, citation entailment, contradiction handling, latency, cost, and adversarial-content handling.
Security evaluation requires a labelled corpus containing genuine injections, benign human instructions, hidden and encoded payloads, tool-poisoning attempts, and extraction failures. Report precision, recall, false-positive rate, false-negative rate, and calibration by risk band. Search benchmarks and security benchmarks should be published separately so strong detection cannot conceal weak retrieval, or vice versa.
Limitations and commitments
- Live-web coverage depends on upstream search availability and index coverage.
- Some pages cannot be extracted because of authentication, robots policies, rendering, regional controls, or network failure.
- Prompt-injection detection is probabilistic and must be evaluated continuously; no detector guarantees safety.
- Source reputation does not prove an individual claim, and first-party sources can be incomplete or self-interested.
- Recent or obscure facts may have only one available source. The system should expose that limitation rather than manufacture agreement.
- Arlong is a retrieval and evidence system, not a substitute for professional judgment in medical, legal, financial, or safety-critical decisions.