Skip to content

AI Verification Pipeline

Business Value

The Labeeb pipeline transforms raw news articles into verified, evidence-backed intelligence through six automated stages. What previously required hours of manual analysis per claim now completes in seconds — enabling fact-checking at a scale that manual processes cannot achieve.


1. Pipeline Overview

How raw content transforms into verified, evidence-backed intelligence.

graph LR
    A[Content Intake] --> B[Quality Control]
    B --> C[Claim Detection]
    C --> D[Evidence Discovery]
    D --> E[Verdict Generation]
    E --> F[Global Delivery]

2. Stage-by-Stage Product Tour

Stage 1: Content Intake

What it does: Monitors 100+ content sources — RSS feeds, APIs, and web pages — and automatically ingests new articles into the platform.

Why it matters: Broad source coverage ensures comprehensive market monitoring without manual effort. New sources deploy through configuration, not code changes.

Metric Value
Sources monitored 100+
Source types RSS, API, Web scraping
Onboarding time Minutes (configuration-based)
Deduplication Automatic at ingestion

A modular adapter system supports multiple source types with template-based configuration profiles. Content is normalized, deduplicated, and persisted to the system of record before entering the processing pipeline.

graph LR
    A[RSS Feeds] --> D[Adapter Layer]
    B[APIs] --> D
    C[Web Pages] --> D
    D --> E[Normalize & Deduplicate]
    E --> F[Store & Queue]

Stage 2: Quality Control

What it does: A multi-tier quality gate filters out low-quality, irrelevant, or problematic content before it enters the AI pipeline.

Why it matters: Prevents wasted processing on non-viable content — filtering approximately 30% of incoming articles and saving proportional downstream costs.

Metric Value
Content filtered ~30%
Assessment speed < 10ms for rule + heuristic tiers
Risk detection Automatic flagging
Quality dimensions Relevance, completeness, language quality

Three progressive evaluation tiers: fast rule-based filtering (< 1ms), heuristic quality assessment (< 10ms), and AI-powered classification for borderline cases. Each tier reduces the volume passed to the next, minimizing compute costs.

graph LR
    A[Incoming Article] --> B[Rule Filter]
    B --> C[Heuristic Assessment]
    C --> D[AI Classification]
    D --> E[Approved for Pipeline]

Stage 3: Claim Detection

What it does: AI identifies which statements in an article are factual claims worthy of verification — separating opinions, questions, and rhetoric from checkworthy statements.

Why it matters: Focuses verification resources on what matters. Humans spend hours reading to find claims; this system identifies them in seconds.

Metric Value
Detection speed Seconds per article
Claim types identified Factual, statistical, attributive
Explainability Confidence scores + routing tags
Reliability Three-tier fallback (primary, batch, heuristic)

A three-tier provider strategy ensures high availability: the primary model provides precise classification, a batch processing mode handles high-volume periods efficiently, and a heuristic fallback ensures claims are always processed even during service disruptions.

graph LR
    A[Article Text] --> B[Sentence Segmentation]
    B --> C[Claim Classification]
    C --> D[Confidence Scoring]
    D --> E[Checkworthy Claims]

Stage 4: Evidence Discovery

What it does: For each identified claim, the system searches multiple sources — internal knowledge bases, external fact-check databases, and the open web — to find corroborating or contradicting evidence.

Why it matters: Comprehensive evidence gathering in under 500ms replaces hours of manual research, and the hybrid search approach finds semantically relevant results that keyword-only search would miss.

Metric Value
Retrieval speed < 500ms
Search approach Hybrid (keyword + semantic)
Sources queried Internal index + external providers
Result quality Reciprocal Rank Fusion scoring

A hybrid retrieval system combines traditional keyword matching (BM25) with semantic vector search (kNN) using Reciprocal Rank Fusion (RRF) to produce high-quality, diverse result sets. External fact-check provider integration adds authoritative third-party evidence.

graph LR
    A[Claim] --> B[Keyword Search]
    A --> C[Semantic Search]
    B --> D[Rank Fusion]
    C --> D
    D --> E[Ranked Evidence]

Stage 5: Verdict Generation

What it does: AI analyzes the relationship between each claim and its retrieved evidence, then generates a transparent verdict with confidence scoring and multilingual explanations.

Why it matters: Produces auditable, evidence-traceable verdicts that analysts and end-users can trust — with full transparency into the reasoning process.

Metric Value
Verdict types True, False, Unclear
Transparency Linked evidence + stance analysis
Languages English and Arabic explanations
Audit trail Complete decision chain logged

A stance detection system classifies each evidence-claim pair as supporting, refuting, or neutral. A weighted voting aggregation then synthesizes multiple stance results into a final verdict with confidence scoring and multilingual explanation generation.

graph LR
    A[Evidence + Claim] --> B[Stance Analysis]
    B --> C[Weighted Voting]
    C --> D[Verdict + Confidence]
    D --> E[Multilingual Explanation]

Stage 6: Global Delivery

What it does: Verified results are indexed and served globally through 300+ edge locations, providing sub-100ms access for analysts, applications, and end-users worldwide.

Why it matters: Instant access to verification results — meeting modern expectations for real-time information delivery regardless of user location.

Metric Value
Edge locations 300+ globally
Response time < 100ms
Delivery methods REST API, streaming, search
Availability 99.5% uptime target

A modular edge architecture handles routing, caching, authentication, and streaming through intelligent service handlers. Results are indexed in both traditional search and vector databases for flexible retrieval patterns.

graph LR
    A[Verified Result] --> B[Search Index]
    A --> C[Vector Index]
    B --> D[Edge Network]
    C --> D
    D --> E[Global Users]

3. End-to-End Impact

Activity Manual Process Automated Pipeline Improvement
Source monitoring Analysts check sources individually 100+ sources monitored continuously Always-on coverage
Quality filtering Editors review each article AI filters in < 10ms 30% waste eliminated
Claim identification Hours of reading per article Seconds per article 100x faster
Evidence gathering Manual database and web searches Hybrid search in < 500ms Comprehensive, instant
Verdict production Written analysis per claim Automated with evidence trails Scalable, consistent
Result delivery Published reports (days) Real-time via edge network Instant global access

4. Strategic Implications

What This Means

The six-stage pipeline creates a verification assembly line that operates continuously without human intervention for standard processing. Each stage builds on the previous one, and the entire chain executes in seconds. This transforms fact-checking from a labor-intensive, per-claim activity into a scalable, automated capability — enabling coverage breadth and speed that manual processes cannot match at any cost.