π Our Scoring Methodology
How we calculate credibility scores for UFO/UAP news and why you can trust our analysis.
The 100-Point System
Every article receives a credibility score from 0-100 based on five weighted factors. This system is designed to be objective, consistent, and transparent. Higher scores indicate stronger evidence and more reliable sources.
The Five Scoring Factors
Source Reliability
The historical accuracy and journalistic standards of the publishing source. Government sources, major news networks, and peer-reviewed publications score highest. Unknown or historically inaccurate sources score lower.
Moderate (60-84): Established specialty publications, regional news
Low (<60): Unknown sources, sites with history of inaccuracy
Evidence Quality
The type and strength of supporting evidence. Physical evidence, official documents, and multi-sensor data score highest. Unsupported claims or anonymous sources score lowest.
Moderate: Photographs, named witness testimony
Weak: Anonymous sources, unsupported claims
Corroboration
Whether multiple independent sources confirm the information. Stories verified by multiple outlets or featuring multiple witnesses score higher than single-source reports.
Moderate: Some corroboration, multiple witnesses
Weak: Single source, exclusive/unverified claims
Official Acknowledgment
Whether government, military, or scientific institutions have acknowledged the phenomenon or information. Official statements and declassified documents significantly boost scores.
Moderate: Acknowledged by agencies but not explained
Weak: No official acknowledgment
Expert Analysis
Whether qualified experts (scientists, former officials, analysts) have weighed in on the claims. Expert endorsement increases credibility; expert debunking decreases it.
Moderate: Some expert commentary, mixed opinions
Weak: No expert analysis or experts dispute claims
Our Analysis Process
1. Aggregation
We continuously monitor news APIs and RSS feeds from trusted sources. When new UFO/UAP content is detected, it enters our analysis queue. We also analyse video β transcribing UFO/UAP footage and interviews so the same scrutiny applies to what's claimed on camera.
2. Source Evaluation
Each source is assigned a base reliability score from our database. Unknown sources receive a default neutral score of 50.
3. Live Context Gathering & Fact-Check
Before any scoring happens, we gather up-to-date context about the story's subject from a live web search, so the analysis isn't based on outdated knowledge. This step:
- searches the web for recent, verified facts and developments on the subject β and for contested or thinly-covered topics, runs additional official, skeptical, and latest-news searches so multiple perspectives are represented;
- reads beyond headlines β it fetches and cleans the full text of the one or two most reliable sources, not just search snippets, for deeper grounding;
- pre-checks the article's own claims, named entities, and any websites it links to β and actually fetches those links to confirm they resolve (official
.gov/.mildomains are treated as authoritative); - flags recent changes β such as an agency being renamed or a new official site launching β so a story isn't penalised for using current terminology that older AI knowledge wouldn't recognise.
Crucially, every fact in the resulting "current-context brief" is tagged with a date and a confidence level. The analysts are told to lean on high-confidence, recent, multiply-sourced facts and to down-weight anything low-confidence or undated β so a rumour never carries the same weight as a documented fact. The brief is cached and reused across articles on the same subject, and is treated as authoritative over the AI's own training data for anything recent.
4. Multi-Agent AI Analysis
Our analysis pipeline has three stages:
Stage 1 β RAG Retrieval: The article is matched against our knowledge graph of linked cases, evidence, and sources, and combined with the live context brief from step 3.
Stage 2 β Five-Agent Debate: Five specialized AI analysts independently evaluate the story: an Evidence Analyst classifies and scores evidence types; a Source Investigator checks source credibility and finds corroborating coverage; a Scientific Skeptic demands empirical proof and identifies alternative explanations; a Historical Archivist compares against known cases and identifies patterns; and a Government Disclosure Analyst tracks official statements and policy context.
Stage 3 β Synthesis: A master Orchestrator reviews all agent outputs, identifies points of agreement and disagreement, runs cross-examinations on disputed areas, and produces a final consensus report with credibility verdict and confidence intervals.
Faster stories are triaged with a lighter single-specialist "Quick" pass; high-profile and contested stories get the full five-agent debate.
5. Score Calculation
Each of the five factors is scored 0-100, then weighted according to the percentages above to produce a final credibility score.
6. Editorial Selection
Scoring a story doesn't automatically publish it. An AI editor reviews everything that's been analysed and ranks it for publishing β weighing credibility, how new it is versus what we already cover, and topical relevance β so the front page stays signal, not noise. The same editor can also work the other way: it independently identifies trending or historically important stories we haven't covered yet, researches them, and feeds them into this very pipeline. Nothing is published automatically unless we've explicitly enabled it for a trusted lane; otherwise a human approves.
7. Report Generation
A detailed report is generated explaining the score breakdown, key findings, and any concerns β with every source we consulted (the original publisher, the live web-context sources, and related cases) listed on the article so you can check our work.
8. Ongoing Freshness
A score is a snapshot of what was known when we analysed it. Because the situation can change, published stories are automatically re-checked and re-scored once their underlying context ages β so an old verdict doesn't quietly go stale. When a re-check is in progress, the article shows a small "score may be outdated β refreshing" note.
Limitations & Disclaimers
No system is perfect. Our scoring methodology is designed to be objective, but it has inherent limitations:
- AI analysis can make errors β we recommend reading full reports
- Source scores are estimates β individual articles may vary in quality
- New evidence can change scores β we update analyses as facts emerge
- A high score doesn't mean "true" β it means strong supporting evidence
- A low score doesn't mean "false" β it means limited verifiable evidence
- Very fresh events may lag β our live context check relies on what the web has already published; for breaking stories we re-run the analysis as coverage catches up
We encourage readers to review our full analyses, check original sources, and form their own conclusions.