• About
  • Advertise
  • Privacy & Policy
  • Contact
Sunday, January 18, 2026
  • Login
  • Home
    • Home – Layout 1
    • Home – Layout 2
    • Home – Layout 3
    • Home – Layout 4
    • Home – Layout 5
    • Home – Layout 6
  • News
    • All
    • Business
    • Politics
    • Science
    • World
    Hillary Clinton in white pantsuit for Trump inauguration

    Hillary Clinton in white pantsuit for Trump inauguration

    Amazon has 143 billion reasons to keep adding more perks to Prime

    Amazon has 143 billion reasons to keep adding more perks to Prime

    Shooting More than 40 Years of New York’s Halloween Parade

    Shooting More than 40 Years of New York’s Halloween Parade

    These Are the 5 Big Tech Stories to Watch in 2017

    These Are the 5 Big Tech Stories to Watch in 2017

    Why Millennials Need to Save Twice as Much as Boomers Did

    Why Millennials Need to Save Twice as Much as Boomers Did

    Doctors take inspiration from online dating to build organ transplant AI

    Doctors take inspiration from online dating to build organ transplant AI

    Trending Tags

    • Trump Inauguration
    • United Stated
    • White House
    • Market Stories
    • Election Results
  • Tech
    • All
    • Apps
    • Gadget
    • Mobile
    • Startup
    The Legend of Zelda: Breath of the Wild gameplay on the Nintendo Switch

    The Legend of Zelda: Breath of the Wild gameplay on the Nintendo Switch

    Shadow Tactics: Blades of the Shogun Review

    Shadow Tactics: Blades of the Shogun Review

    macOS Sierra review: Mac users get a modest update this year

    macOS Sierra review: Mac users get a modest update this year

    Hands on: Samsung Galaxy A5 2017 review

    Hands on: Samsung Galaxy A5 2017 review

    The Last Guardian Playstation 4 Game review

    The Last Guardian Playstation 4 Game review

    These Are the 5 Big Tech Stories to Watch in 2017

    These Are the 5 Big Tech Stories to Watch in 2017

    Trending Tags

    • Nintendo Switch
    • CES 2017
    • Playstation 4 Pro
    • Mark Zuckerberg
  • Entertainment
    • All
    • Gaming
    • Movie
    • Music
    • Sports
    The Legend of Zelda: Breath of the Wild gameplay on the Nintendo Switch

    The Legend of Zelda: Breath of the Wild gameplay on the Nintendo Switch

    macOS Sierra review: Mac users get a modest update this year

    macOS Sierra review: Mac users get a modest update this year

    Hands on: Samsung Galaxy A5 2017 review

    Hands on: Samsung Galaxy A5 2017 review

    Heroes of the Storm Global Championship 2017 starts tomorrow, here’s what you need to know

    Heroes of the Storm Global Championship 2017 starts tomorrow, here’s what you need to know

    Harnessing the power of VR with Power Rangers and Snapdragon 835

    Harnessing the power of VR with Power Rangers and Snapdragon 835

    So you want to be a startup investor? Here are things you should know

    So you want to be a startup investor? Here are things you should know

  • Lifestyle
    • All
    • Fashion
    • Food
    • Health
    • Travel
    Shooting More than 40 Years of New York’s Halloween Parade

    Shooting More than 40 Years of New York’s Halloween Parade

    Heroes of the Storm Global Championship 2017 starts tomorrow, here’s what you need to know

    Heroes of the Storm Global Championship 2017 starts tomorrow, here’s what you need to know

    Why Millennials Need to Save Twice as Much as Boomers Did

    Why Millennials Need to Save Twice as Much as Boomers Did

    Doctors take inspiration from online dating to build organ transplant AI

    Doctors take inspiration from online dating to build organ transplant AI

    How couples can solve lighting disagreements for good

    How couples can solve lighting disagreements for good

    Ducati launch: Lorenzo and Dovizioso’s Desmosedici

    Ducati launch: Lorenzo and Dovizioso’s Desmosedici

    Trending Tags

    • Golden Globes
    • Game of Thrones
    • MotoGP 2017
    • eSports
    • Fashion Week
  • Review
    The Legend of Zelda: Breath of the Wild gameplay on the Nintendo Switch

    The Legend of Zelda: Breath of the Wild gameplay on the Nintendo Switch

    Shadow Tactics: Blades of the Shogun Review

    Shadow Tactics: Blades of the Shogun Review

    macOS Sierra review: Mac users get a modest update this year

    macOS Sierra review: Mac users get a modest update this year

    Hands on: Samsung Galaxy A5 2017 review

    Hands on: Samsung Galaxy A5 2017 review

    The Last Guardian Playstation 4 Game review

    The Last Guardian Playstation 4 Game review

    Intel Core i7-7700K ‘Kaby Lake’ review

    Intel Core i7-7700K ‘Kaby Lake’ review

No Result
View All Result
Ai News
Advertisement
  • Home
    • Home – Layout 1
    • Home – Layout 2
    • Home – Layout 3
    • Home – Layout 4
    • Home – Layout 5
    • Home – Layout 6
  • News
    • All
    • Business
    • Politics
    • Science
    • World
    Hillary Clinton in white pantsuit for Trump inauguration

    Hillary Clinton in white pantsuit for Trump inauguration

    Amazon has 143 billion reasons to keep adding more perks to Prime

    Amazon has 143 billion reasons to keep adding more perks to Prime

    Shooting More than 40 Years of New York’s Halloween Parade

    Shooting More than 40 Years of New York’s Halloween Parade

    These Are the 5 Big Tech Stories to Watch in 2017

    These Are the 5 Big Tech Stories to Watch in 2017

    Why Millennials Need to Save Twice as Much as Boomers Did

    Why Millennials Need to Save Twice as Much as Boomers Did

    Doctors take inspiration from online dating to build organ transplant AI

    Doctors take inspiration from online dating to build organ transplant AI

    Trending Tags

    • Trump Inauguration
    • United Stated
    • White House
    • Market Stories
    • Election Results
  • Tech
    • All
    • Apps
    • Gadget
    • Mobile
    • Startup
    The Legend of Zelda: Breath of the Wild gameplay on the Nintendo Switch

    The Legend of Zelda: Breath of the Wild gameplay on the Nintendo Switch

    Shadow Tactics: Blades of the Shogun Review

    Shadow Tactics: Blades of the Shogun Review

    macOS Sierra review: Mac users get a modest update this year

    macOS Sierra review: Mac users get a modest update this year

    Hands on: Samsung Galaxy A5 2017 review

    Hands on: Samsung Galaxy A5 2017 review

    The Last Guardian Playstation 4 Game review

    The Last Guardian Playstation 4 Game review

    These Are the 5 Big Tech Stories to Watch in 2017

    These Are the 5 Big Tech Stories to Watch in 2017

    Trending Tags

    • Nintendo Switch
    • CES 2017
    • Playstation 4 Pro
    • Mark Zuckerberg
  • Entertainment
    • All
    • Gaming
    • Movie
    • Music
    • Sports
    The Legend of Zelda: Breath of the Wild gameplay on the Nintendo Switch

    The Legend of Zelda: Breath of the Wild gameplay on the Nintendo Switch

    macOS Sierra review: Mac users get a modest update this year

    macOS Sierra review: Mac users get a modest update this year

    Hands on: Samsung Galaxy A5 2017 review

    Hands on: Samsung Galaxy A5 2017 review

    Heroes of the Storm Global Championship 2017 starts tomorrow, here’s what you need to know

    Heroes of the Storm Global Championship 2017 starts tomorrow, here’s what you need to know

    Harnessing the power of VR with Power Rangers and Snapdragon 835

    Harnessing the power of VR with Power Rangers and Snapdragon 835

    So you want to be a startup investor? Here are things you should know

    So you want to be a startup investor? Here are things you should know

  • Lifestyle
    • All
    • Fashion
    • Food
    • Health
    • Travel
    Shooting More than 40 Years of New York’s Halloween Parade

    Shooting More than 40 Years of New York’s Halloween Parade

    Heroes of the Storm Global Championship 2017 starts tomorrow, here’s what you need to know

    Heroes of the Storm Global Championship 2017 starts tomorrow, here’s what you need to know

    Why Millennials Need to Save Twice as Much as Boomers Did

    Why Millennials Need to Save Twice as Much as Boomers Did

    Doctors take inspiration from online dating to build organ transplant AI

    Doctors take inspiration from online dating to build organ transplant AI

    How couples can solve lighting disagreements for good

    How couples can solve lighting disagreements for good

    Ducati launch: Lorenzo and Dovizioso’s Desmosedici

    Ducati launch: Lorenzo and Dovizioso’s Desmosedici

    Trending Tags

    • Golden Globes
    • Game of Thrones
    • MotoGP 2017
    • eSports
    • Fashion Week
  • Review
    The Legend of Zelda: Breath of the Wild gameplay on the Nintendo Switch

    The Legend of Zelda: Breath of the Wild gameplay on the Nintendo Switch

    Shadow Tactics: Blades of the Shogun Review

    Shadow Tactics: Blades of the Shogun Review

    macOS Sierra review: Mac users get a modest update this year

    macOS Sierra review: Mac users get a modest update this year

    Hands on: Samsung Galaxy A5 2017 review

    Hands on: Samsung Galaxy A5 2017 review

    The Last Guardian Playstation 4 Game review

    The Last Guardian Playstation 4 Game review

    Intel Core i7-7700K ‘Kaby Lake’ review

    Intel Core i7-7700K ‘Kaby Lake’ review

No Result
View All Result
Ai News
No Result
View All Result
Home Machine Learning

A Geometric Method to Spot Hallucinations Without an LLM Judge

AiNEWS2025 by AiNEWS2025
2026-01-18
in Machine Learning
0
A Geometric Method to Spot Hallucinations Without an LLM Judge
0
SHARES
0
VIEWS
Share on FacebookShare on Twitter


of birds in flight.

There’s no leader. No central command. Each bird aligns with its neighbors—matching direction, adjusting speed, maintaining coherence through purely local coordination. The result is global order emerging from local consistency.

Now imagine one bird flying with the same conviction as the others. Its wingbeats are confident. Its speed is correct. But its direction doesn’t match its neighbors. It’s the red bird.

It’s not lost. It’s not hesitating. It simply doesn’t belong to the flock.

Hallucinations in LLMs are red birds.

The problem we’re actually trying to solve

LLMs generate fluent, confident text that may contain fabricated information. They invent legal cases that don’t exist. They cite papers that were never written. They state facts with the same tone whether those facts are true or completely made up.

The standard approach to detecting this is to ask another language model to check the output. LLM-as-judge. You can see the problem immediately: we’re using a system that hallucinates to detect hallucinations. It’s like asking someone who can’t distinguish colors to sort paint samples. They’ll give you an answer. It might even be right sometimes. But they’re not actually seeing what you need them to see.

The question we asked was different: can we detect hallucinations from the geometric structure of the text itself, without needing another language model’s opinion?

What embeddings actually do

Before getting to the detection method, I want to step back and establish what we’re working with.

When you feed text into a sentence encoder, you get back a vector—a point in high-dimensional space. Texts that are semantically similar land near each other. Texts that are unrelated land far apart. This is what contrastive training optimizes for. But there’s a more subtle tructure than just “similar things are close.”

Consider what happens when you embed a question and its answer. The question lands somewhere in this embeddings space. The answer lands somewhere else. The vector connecting them—what we call the displacement—points in a particular direction. We have a vector: a magnitude and an angle.

We also observed that for grounded responses within a specific domain, these displacement vectors point in consistent directions. We have found something in common: angles.

If you ask five similar questions and get five grounded answers, the displacements from question to answer will be roughly parallel. Not identical—the magnitudes vary, the exact angles differ slightly—but the overall direction is consistent.

When a model hallucinates, something different happens. The response still lands somewhere in embedding space. It’s still fluent. It still sounds like an answer. But the displacement doesn’t follow the local pattern. It points elsewhere. A vector with a totally different angle.

The red bird is flying confidently. But not with the flock. Flies in the opposite direction with an angle totally different from the rest of the birds.

Displacement Consistency (DC)

We formalize this as Displacement Consistency (DC). The idea is simple:

  1. Build a reference set of grounded question-answer pairs from your domain
  2. For a new question-answer pair, find the neighboring questions in the reference set
  3. Compute the mean displacement direction of those neighbors
  4. Measure how well the new displacement aligns with that mean direction

Grounded responses align well. Hallucinated responses don’t. That’s it. One cosine similarity. No source documents needed at inference time. No multiple generations. No model internals.

And it works remarkably well. Across five architecturally distinct embedding models, across multiple hallucination benchmarks including HaluEval and TruthfulQA, DC achieves near-perfect discrimination. The distributions barely overlap.

The catch: domain locality

We tested DC across five embedding models chosen to span architectural diversity: MPNet-based contrastive fine-tuning (all-mpnet-base-v2), weakly-supervised pre-training (E5-large-v2), instruction-tuned training with hard negatives (BGE-large-en-v1.5), encoder-decoder adaptation (GTR-T5-large), and efficient long-context architectures (nomic-embed-text-v1.5). If DC only worked with one architecture, it might be an artifact of that specific model. Consistent results across architecturally distinct models would suggest the structure is fundamental.

The results were consistent. DC achieved AUROC of 1.0 across all five models on our synthetic benchmark. But synthetic benchmarks can be misleading—perhaps domain-shuffled responses are simply too easy to detect.

So we validated on established hallucination datasets: HaluEval-QA, which contains LLM-generated hallucinations specifically designed to be subtle; HaluEval-Dialogue, with responses that deviate from conversation context; and TruthfulQA, which tests common misconceptions that humans frequently believe.

DC maintained perfect discrimination on all of them. Zero degradation from synthetic to realistic benchmarks.

For comparison, ratio-based methods that measure where responses land relative to queries (rather than the direction they move) achieved AUROC around 0.70–0.81. The gap—approximately 0.20 absolute AUROC—is substantial and consistent across all models tested.

The score distributions tell the story visually. Grounded responses cluster tightly at high DC values (around 0.9). Hallucinated responses spread at lower values (around 0.3). The distributions barely overlap.

DC achieves perfect detection within a narrow domain. But if you try to use a reference set from one domain to detect hallucinations in another domain, performance drops to random—AUROC around 0.50. This is telling us something fundamental about how embeddings encode grounding. It is equivalent to see different flocks in the sky: every flock will have a different direction.

For LLMs, the easiest way to understand this is through the image of what in geometry is called a “fiber bundle”.

Figure 1. Geometric fiber bundle. Image by author.

The surface in Figure 1 is the base manifold representing all possible questions. At each point on this surface, there’s a fiber: a line pointing in the direction that grounded responses move. Within any local region of the surface (one specific domain), all the fibers point roughly the same way. That’s why DC works so well locally.

But globally, across different regions, the fibers point in different directions. The “grounded direction” for legal questions is different from the “grounded direction” for medical questions. There’s no single global pattern. Only local coherence.

Now look at the following video. Birds flight paths connecting Europe and Africa. We can see the fiber bundles. Different birds (medium/large small, insects) have different directions.

Video Copyright from https://www.arcgis.com/. Use according 2.2 Grant of Noncommercial Use of Services. Noncommercial Use may include teaching, classroom use, scholarship, and/or research, subject to the fair use rights enumerated in sections 107 and 108 of the Copyright Act (Title 17 of the United States Code).

In differential geometry, this structure is called local triviality without global triviality. Each patch of the manifold looks simple and consistent internally. But the patches can’t be stitched together into one global coordinate system.

This has a noticeable implication:

grounding is not a universal geometric property

There’s no single “truthfulness direction” in embedding space. Each domain—each type of task, each LLM—develops its own displacement pattern during training. The patterns are real and detectable, but they’re domain-specific. Birds do not migrate in the same direction.

What this means practically

For deployment, the domain-locality finding means you need a small calibration set (around 100 examples) matched to your specific use case. A legal Q&A system needs legal examples. A medical chatbot needs medical examples. This is a one-time upfront cost—the calibration happens offline—but it can’t be skipped.

For understanding embeddings, the finding suggests these models encode richer structure than we typically assume. They’re not just learning “similarity.” They’re learning domain-specific mappings whose disruption reliably signals hallucination.

The red bird doesn’t d

The hallucinated response has no marker that says “I’m fabricated.” It’s fluent. It’s confident. It looks exactly like a grounded response on every surface-level metric.

But it doesn’t move with the flock. And now we can measure that.

The geometry has been there all along, implicit in how contrastive training shapes embedding space. We’re just learning to read it.


Notes:

You can find the complete paper at https://cert-framework.com/docs/research/dc-paper.

If you have any questions about the discussed topics, feel free to contact me at [email protected]

Source link

#Geometric #Method #Spot #Hallucinations #LLM #Judge

Tags: Ai Hallucinationartificial intelligenceLlmLlm Applicationsmachine learning
Previous Post

Managers on alert for “launch fever” as pressure builds for NASA’s Moon mission

AiNEWS2025

AiNEWS2025

Stay Connected test

  • 23.9k Followers
  • 99 Subscribers
  • Trending
  • Comments
  • Latest
A tiny new open source AI model performs as well as powerful big ones

A tiny new open source AI model performs as well as powerful big ones

0
Water Cooler Small Talk: The Birthday Paradox 🎂🎉 | by Maria Mouschoutzi, PhD | Sep, 2024

Water Cooler Small Talk: The Birthday Paradox 🎂🎉 | by Maria Mouschoutzi, PhD | Sep, 2024

0
Ghost of Yōtei: The acclaimed Ghost of Tsushima is getting a sequel

Ghost of Yōtei: The acclaimed Ghost of Tsushima is getting a sequel

0
Best Headphones for Working Out (2024): Bose, Shokz, JLab

Best Headphones for Working Out (2024): Bose, Shokz, JLab

0
A Geometric Method to Spot Hallucinations Without an LLM Judge

A Geometric Method to Spot Hallucinations Without an LLM Judge

2026-01-18
Managers on alert for “launch fever” as pressure builds for NASA’s Moon mission

Managers on alert for “launch fever” as pressure builds for NASA’s Moon mission

2026-01-18
The Setapp Mobile iOS store is shutting down on February 16th

The Setapp Mobile iOS store is shutting down on February 16th

2026-01-18
The Workers Building Labubus Are Allegedly Being Horribly Exploited

The Workers Building Labubus Are Allegedly Being Horribly Exploited

2026-01-18

Recent News

A Geometric Method to Spot Hallucinations Without an LLM Judge

A Geometric Method to Spot Hallucinations Without an LLM Judge

2026-01-18
Managers on alert for “launch fever” as pressure builds for NASA’s Moon mission

Managers on alert for “launch fever” as pressure builds for NASA’s Moon mission

2026-01-18
The Setapp Mobile iOS store is shutting down on February 16th

The Setapp Mobile iOS store is shutting down on February 16th

2026-01-18
The Workers Building Labubus Are Allegedly Being Horribly Exploited

The Workers Building Labubus Are Allegedly Being Horribly Exploited

2026-01-18
Footer logo

We bring you the best Premium WordPress Themes that perfect for news, magazine, personal blog, etc. Check our landing page for details.

Follow Us

Browse by Category

  • AI & Cloud Computing
  • AI & Cybersecurity
  • AI & Sentiment Analysis
  • AI Applications
  • AI Ethics
  • AI Future Predictions
  • AI in Education
  • AI in Fintech
  • AI in Gaming
  • AI in Healthcare
  • AI in Startups
  • AI Innovations
  • AI News
  • AI Research
  • AI Tools & Automation
  • Apps
  • AR/VR & AI
  • Business
  • Deep Learning
  • Emerging Technologies
  • Entertainment
  • Fashion
  • Food
  • Gadget
  • Gaming
  • Health
  • Lifestyle
  • Machine Learning
  • Mobile
  • Movie
  • Music
  • News
  • Politics
  • Review
  • Robotics & Smart Systems
  • Science
  • Sports
  • Startup
  • Tech
  • Travel
  • World

Recent News

A Geometric Method to Spot Hallucinations Without an LLM Judge

A Geometric Method to Spot Hallucinations Without an LLM Judge

2026-01-18
Managers on alert for “launch fever” as pressure builds for NASA’s Moon mission

Managers on alert for “launch fever” as pressure builds for NASA’s Moon mission

2026-01-18
  • About
  • Advertise
  • Privacy & Policy
  • Contact

© 2026 JNews - Premium WordPress news & magazine theme by Jegtheme.

Welcome Back!

Login to your account below

Forgotten Password?

Retrieve your password

Please enter your username or email address to reset your password.

Log In
No Result
View All Result

© 2026 JNews - Premium WordPress news & magazine theme by Jegtheme.