Close Menu
journearn.comjournearn.com
  • Home
  • Apps
  • Business
  • Make Money Online
  • Money Saving
  • Finance
  • Food
  • Investment
  • Travel
Facebook X (Twitter) Instagram
journearn.comjournearn.com
Facebook Instagram Pinterest Vimeo
  • Home
  • Apps

    IP Ownership Architecture for India ODC Engagements

    July 24, 2026

    Engineering Capability Guide for ISVs

    July 22, 2026

    How AI Solves Supply Chain Risk Monitoring? 8 Use Cases in 2026

    July 20, 2026

    Top 10 Apple TV App Development Companies (2026 Ranking)

    July 18, 2026

    The ROI Case for Digitizing Your Yard in 2026

    July 16, 2026
  • Business

    7 Key Dates for the First Day of Filing Taxes 2025

    August 3, 2026

    As Americans’ Debt Grows So Does Bankruptcy –

    August 2, 2026

    How to Track Your Brand’s AI Visiblity in 2026 

    August 2, 2026

    50+ Sales Email Templates for Every Stage of the Pipeline

    August 1, 2026

    I Tested 8 Best Employee Monitoring Software for 2026: See My Review

    July 31, 2026
  • Make Money Online

    Amazon Got $600 Million of Your Tariff Money Back. 6 Things Every Shopper Should Know

    August 2, 2026

    How One Couple Erased $40,000 of Debt in 18 Months (Without Eating Ramen)

    July 31, 2026

    271. “He hid $30K of debt a month before our wedding”

    July 29, 2026

    How Trust, Emotions and Chemistry Are Reshaping the American Workforce in 2026

    July 28, 2026

    Buc-ee’s Pay Chart Goes Viral. See How Much Employees Can Earn

    July 26, 2026
  • Money Saving

    Received a Social Security Overpayment Notice? Take These Steps Before You Pay

    August 3, 2026

    Stock news for investors: First Quantum profit jumps, Intact earnings fall

    August 2, 2026

    What Is the Best Credit Card for Students Studying Abroad? A 3-Factor Decision Guide

    August 1, 2026

    Why families buy alloy wheels at the wrong time (and how to avoid overpaying)

    July 31, 2026

    WIN! DenTek Fun Summer Smile prize bundle

    July 30, 2026
  • Finance

    One Late Payment Can Stay on Your Credit Report for Seven Years

    August 2, 2026

    Once CPP disability benefits and an annuity following a car crash end at 65, should Anita switch to a TFSA and RRSP?

    August 1, 2026

    Switching Banks Without Missing a Single Bill Payment

    July 30, 2026

    ProjectionLab Review 2026: Best DIY Retirement Planning Tool

    July 29, 2026

    Your Credit Utilization Should Stay Below 30%

    July 27, 2026
  • Food

    Homemade Graham Crackers – Sally’s Baking

    August 3, 2026

    Lemon Delicious Pudding – Extra Saucy!

    August 2, 2026

    Eater World’s Fare Was a Global Celebration of Food and Soccer

    July 31, 2026

    Panisses (Chickpea Fries) – Cookie and Kate

    July 30, 2026

    Cherry Clafoutis | Skinnytaste

    July 29, 2026
  • Investment

    The Pieces of AGI Are Falling Into Place

    August 3, 2026

    Silver Sector M&A Hits US$14.3 Billion As Miners Hunt for Growth

    August 2, 2026

    New Windows Update Makes Dell PCs Just Shutdown

    August 1, 2026

    Can Cost Segregation Studies Help If I Bought the Property Years Ago?

    July 31, 2026

    The SEC’s Proposal for Semiannual Reporting

    July 30, 2026
  • Travel

    Anantara Layan Phuket: The Luxury Resort That Let Me Feel Like Myself Again

    August 2, 2026

    Orlando Staycation Guide: for Florida & Georgia Residents

    August 1, 2026

    Your First Week Home Decides How Long Your Trip Lasts

    August 1, 2026

    Is Waldorf Astoria Costa Rica Punta Cacique Worth It? Our Family Stay Review

    July 29, 2026

    Visiting The Gold Coast Australia: 10 BEST Things To Do

    July 28, 2026
journearn.comjournearn.com
Home»Investment»Chart of the Week: AI Is a Black Box
Investment

Chart of the Week: AI Is a Black Box

info@journearn.comBy info@journearn.comJune 18, 2026No Comments4 Mins Read
Facebook Twitter Pinterest LinkedIn Tumblr WhatsApp Telegram Email
Chart of the Week: AI Is a Black Box
Share
Facebook Twitter LinkedIn Pinterest Email


A strange thing happened last week.

Anthropic was forced to take its newest AI models offline only days after releasing them.

The company’s new Fable 5 and Mythos 5 systems were designed to be some of the most powerful AI models ever released. But shortly after launch, researchers discovered ways to get around some of the models’ built-in safety measures.

Government officials soon got involved as fears spread that these systems could become powerful cybersecurity weapons in the wrong hands.

Maybe those concerns were justified, and maybe they weren’t.

But to me, they raise an obvious question that not enough people are asking.

How would anyone know?

What’s Inside the Box?

Modern AI systems aren’t like traditional software.

Engineers don’t sit down and write lines of code telling them exactly how to reason through a problem.

Instead, researchers train these systems and then observe their behavior.

The result is what many researchers call a black box.

We can see what goes in, and we can see what comes out.

But what happens in between is often much harder to explain.

That’s why companies like Anthropic spend so much time studying AI interpretability, or the science of understanding how these systems arrive at their conclusions.

And that brings us to this week’s chart.

Because a group of researchers recently performed a strange experiment.

They secretly modified an AI model’s internal state. Then they asked whether the model could detect that something had changed.

AI interpretability experiment

Image: Uzay Macar and Li Yang

This chart might look complicated, but the basic idea is simple.

Researchers injected information directly into an AI model’s internal processing, then tested whether it could tell the difference between those injections and its normal thought process.

The chart compares three versions of the same model.

The first is the Base model, the raw AI system before it receives additional training.

The second is the Instruct model, which was trained to behave more like the helpful AI assistants most people interact with today.

The third is an Abliterated version of the model, where some of the refusal and safety behaviors were removed.

The blue line shows how often the model correctly detected a real change, while the orange line shows how often it falsely claimed that something changed when nothing had actually happened.

And the results are surprising.

The Base model performed poorly. When researchers secretly altered its internal processing, it often couldn’t tell the difference between a real change and a false alarm.

But the Instruct model performed much better.

Somewhere during the additional training process, the model appears to have developed an ability to recognize when something unusual had happened inside its own processing.

And in several cases, the Abliterated model performed even better still.

In other words, removing some of the AI’s safety and refusal behaviors actually improved the model’s ability to detect what was going on inside it.

That doesn’t mean the model became conscious or self-aware.

You can compare it to a computer server that detects when someone has tampered with its memory. The server isn’t aware of anything, but it can still recognize when something unusual has happened.

Researchers believe something similar happened here.

More importantly, they think capabilities like this could eventually help us better understand what’s happening inside advanced AI systems.

After all, these models have access to information that remains largely hidden from the people studying them.

Which means one way researchers could eventually learn more about advanced AI systems is by asking the systems themselves.

That might seem counterintuitive.

But it would give researchers something they’ve never really had before.

A window into what’s happening inside the model itself.

Here’s My Take

The primary goal of the AI industry has been to build more capable models.

But another challenge is gaining urgency.

Understanding them.

The controversy surrounding Anthropic’s latest models shows why we need to get a handle on this issue sooner than later.

Because it’s one thing to build a powerful AI system. It’s something else entirely to create a new form of intelligence yet only partially understand how it works.

So here’s my question to you:

If future AI systems become too complex for humans to fully understand on their own, would you trust AI to help explain what’s happening inside other AI models?

Or does that sound like asking the fox to guard the henhouse?

I’d love to hear what you think.

Let me know at dailydisruptor@banyanhill.com.

We won’t reveal your full name in the event we publish a response, so feel free to share your honest opinion.

Regards,

Ian King's Signature
Ian King
Chief Strategist, Banyan Hill Publishing





Source link

AI Anthropic
Share. Facebook Twitter Pinterest LinkedIn Tumblr Email
info
info@journearn.com
  • Website

Related Posts

The Pieces of AGI Are Falling Into Place

August 3, 2026

Silver Sector M&A Hits US$14.3 Billion As Miners Hunt for Growth

August 2, 2026

New Windows Update Makes Dell PCs Just Shutdown

August 1, 2026

Can Cost Segregation Studies Help If I Bought the Property Years Ago?

July 31, 2026

The SEC’s Proposal for Semiannual Reporting

July 30, 2026

AI Just Broke Free – Banyan Hill Publishing

July 29, 2026
Add A Comment
Leave A Reply Cancel Reply

  • Facebook
  • Twitter
  • Instagram
  • Pinterest
Don't Miss

The Pieces of AGI Are Falling Into Place

Received a Social Security Overpayment Notice? Take These Steps Before You Pay

7 Key Dates for the First Day of Filing Taxes 2025

Homemade Graham Crackers – Sally’s Baking

About Us

Welcome to Journearn.com – your trusted guide on the journey to earning smarter, saving better, and building a more financially secure future. At Journearn, we believe that financial knowledge should be accessible to everyone.

Quicklinks
  • Business
  • Food
  • Make Money Online
  • Money Saving
  • Travel
Useful Links
  • About Us
  • Contact Us
  • Disclaimer
  • Privacy Policy
  • Terms and Conditions
Popular Posts

The Pieces of AGI Are Falling Into Place

August 3, 2026

Received a Social Security Overpayment Notice? Take These Steps Before You Pay

August 3, 2026
© 2026 Designed by journearn.All Right Reserved

Type above and press Enter to search. Press Esc to cancel.