AI Verification Pipeline¶
Business Value
The Labeeb pipeline transforms raw news articles into verified, evidence-backed intelligence through six automated stages. What previously required hours of manual analysis per claim now completes in seconds — enabling fact-checking at a scale that manual processes cannot achieve.
1. Pipeline Overview¶
How raw content transforms into verified, evidence-backed intelligence.
graph LR
A[Content Intake] --> B[Quality Control]
B --> C[Claim Detection]
C --> D[Evidence Discovery]
D --> E[Verdict Generation]
E --> F[Global Delivery]
2. Stage-by-Stage Product Tour¶
Stage 1: Content Intake¶
What it does: Monitors 100+ content sources — RSS feeds, APIs, and web pages — and automatically ingests new articles into the platform.
Why it matters: Broad source coverage ensures comprehensive market monitoring without manual effort. New sources deploy through configuration, not code changes.
| Metric | Value |
|---|---|
| Sources monitored | 100+ |
| Source types | RSS, API, Web scraping |
| Onboarding time | Minutes (configuration-based) |
| Deduplication | Automatic at ingestion |
A modular adapter system supports multiple source types with template-based configuration profiles. Content is normalized, deduplicated, and persisted to the system of record before entering the processing pipeline.
graph LR
A[RSS Feeds] --> D[Adapter Layer]
B[APIs] --> D
C[Web Pages] --> D
D --> E[Normalize & Deduplicate]
E --> F[Store & Queue]
Stage 2: Quality Control¶
What it does: A multi-tier quality gate filters out low-quality, irrelevant, or problematic content before it enters the AI pipeline.
Why it matters: Prevents wasted processing on non-viable content — filtering approximately 30% of incoming articles and saving proportional downstream costs.
| Metric | Value |
|---|---|
| Content filtered | ~30% |
| Assessment speed | < 10ms for rule + heuristic tiers |
| Risk detection | Automatic flagging |
| Quality dimensions | Relevance, completeness, language quality |
Three progressive evaluation tiers: fast rule-based filtering (< 1ms), heuristic quality assessment (< 10ms), and AI-powered classification for borderline cases. Each tier reduces the volume passed to the next, minimizing compute costs.
graph LR
A[Incoming Article] --> B[Rule Filter]
B --> C[Heuristic Assessment]
C --> D[AI Classification]
D --> E[Approved for Pipeline]
Stage 3: Claim Detection¶
What it does: AI identifies which statements in an article are factual claims worthy of verification — separating opinions, questions, and rhetoric from checkworthy statements.
Why it matters: Focuses verification resources on what matters. Humans spend hours reading to find claims; this system identifies them in seconds.
| Metric | Value |
|---|---|
| Detection speed | Seconds per article |
| Claim types identified | Factual, statistical, attributive |
| Explainability | Confidence scores + routing tags |
| Reliability | Three-tier fallback (primary, batch, heuristic) |
A three-tier provider strategy ensures high availability: the primary model provides precise classification, a batch processing mode handles high-volume periods efficiently, and a heuristic fallback ensures claims are always processed even during service disruptions.
graph LR
A[Article Text] --> B[Sentence Segmentation]
B --> C[Claim Classification]
C --> D[Confidence Scoring]
D --> E[Checkworthy Claims]
Stage 4: Evidence Discovery¶
What it does: For each identified claim, the system searches multiple sources — internal knowledge bases, external fact-check databases, and the open web — to find corroborating or contradicting evidence.
Why it matters: Comprehensive evidence gathering in under 500ms replaces hours of manual research, and the hybrid search approach finds semantically relevant results that keyword-only search would miss.
| Metric | Value |
|---|---|
| Retrieval speed | < 500ms |
| Search approach | Hybrid (keyword + semantic) |
| Sources queried | Internal index + external providers |
| Result quality | Reciprocal Rank Fusion scoring |
A hybrid retrieval system combines traditional keyword matching (BM25) with semantic vector search (kNN) using Reciprocal Rank Fusion (RRF) to produce high-quality, diverse result sets. External fact-check provider integration adds authoritative third-party evidence.
graph LR
A[Claim] --> B[Keyword Search]
A --> C[Semantic Search]
B --> D[Rank Fusion]
C --> D
D --> E[Ranked Evidence]
Stage 5: Verdict Generation¶
What it does: AI analyzes the relationship between each claim and its retrieved evidence, then generates a transparent verdict with confidence scoring and multilingual explanations.
Why it matters: Produces auditable, evidence-traceable verdicts that analysts and end-users can trust — with full transparency into the reasoning process.
| Metric | Value |
|---|---|
| Verdict types | True, False, Unclear |
| Transparency | Linked evidence + stance analysis |
| Languages | English and Arabic explanations |
| Audit trail | Complete decision chain logged |
A stance detection system classifies each evidence-claim pair as supporting, refuting, or neutral. A weighted voting aggregation then synthesizes multiple stance results into a final verdict with confidence scoring and multilingual explanation generation.
graph LR
A[Evidence + Claim] --> B[Stance Analysis]
B --> C[Weighted Voting]
C --> D[Verdict + Confidence]
D --> E[Multilingual Explanation]
Stage 6: Global Delivery¶
What it does: Verified results are indexed and served globally through 300+ edge locations, providing sub-100ms access for analysts, applications, and end-users worldwide.
Why it matters: Instant access to verification results — meeting modern expectations for real-time information delivery regardless of user location.
| Metric | Value |
|---|---|
| Edge locations | 300+ globally |
| Response time | < 100ms |
| Delivery methods | REST API, streaming, search |
| Availability | 99.5% uptime target |
A modular edge architecture handles routing, caching, authentication, and streaming through intelligent service handlers. Results are indexed in both traditional search and vector databases for flexible retrieval patterns.
graph LR
A[Verified Result] --> B[Search Index]
A --> C[Vector Index]
B --> D[Edge Network]
C --> D
D --> E[Global Users]
3. End-to-End Impact¶
| Activity | Manual Process | Automated Pipeline | Improvement |
|---|---|---|---|
| Source monitoring | Analysts check sources individually | 100+ sources monitored continuously | Always-on coverage |
| Quality filtering | Editors review each article | AI filters in < 10ms | 30% waste eliminated |
| Claim identification | Hours of reading per article | Seconds per article | 100x faster |
| Evidence gathering | Manual database and web searches | Hybrid search in < 500ms | Comprehensive, instant |
| Verdict production | Written analysis per claim | Automated with evidence trails | Scalable, consistent |
| Result delivery | Published reports (days) | Real-time via edge network | Instant global access |
4. Strategic Implications¶
What This Means
The six-stage pipeline creates a verification assembly line that operates continuously without human intervention for standard processing. Each stage builds on the previous one, and the entire chain executes in seconds. This transforms fact-checking from a labor-intensive, per-claim activity into a scalable, automated capability — enabling coverage breadth and speed that manual processes cannot match at any cost.