<?xml version="1.0" encoding="utf-8" standalone="yes"?><rss version="2.0" xmlns:atom="http://www.w3.org/2005/Atom"><channel><title>Tobias's Blog</title><link>https://bertobi.github.io/blog/</link><description>Recent content on Tobias's Blog</description><generator>Hugo</generator><language>en-us</language><lastBuildDate>Sat, 13 Jun 2026 00:00:00 +0000</lastBuildDate><atom:link href="https://bertobi.github.io/blog/index.xml" rel="self" type="application/rss+xml"/><item><title>ARENA 8.0 Hackathon - NLAs all the way down</title><link>https://bertobi.github.io/blog/posts/2026-06-13-arena-hackathon-nlas-all-the-way-down/</link><pubDate>Sat, 13 Jun 2026 00:00:00 +0000</pubDate><guid>https://bertobi.github.io/blog/posts/2026-06-13-arena-hackathon-nlas-all-the-way-down/</guid><description>&lt;h1 id="nlas-all-the-way-down-a-verbalizers-self-readout-doesnt-reveal-when-its-making-things-up"&gt;NLAs all the way down: a verbalizer&amp;rsquo;s self-readout doesn&amp;rsquo;t reveal when it&amp;rsquo;s making things up&lt;/h1&gt;
&lt;p&gt;&lt;em&gt;ARENA hackathon, June 2026.&lt;/em&gt;&lt;/p&gt;
&lt;h2 id="tldr"&gt;TL;DR&lt;/h2&gt;
&lt;p&gt;A &lt;a href="https://transformer-circuits.pub/2026/nla/index.html"&gt;natural language autoencoder (NLA)&lt;/a&gt; is a tool that translates a model&amp;rsquo;s internal activations into English. We (Me and Claude), used the NLA to explain its own internal state. When the verbalizer made up a fact, we read its own activations and passed them through the same NLA to see whether the readout would give signs of it being a fabrication. It didn&amp;rsquo;t. A confabulating state and a faithful one produce the same confident description, with no hedging and no sign that anything was invented. Our reading: confabulation is confident, the model isn&amp;rsquo;t internally flagging doubt, so a self-readout has nothing to surface.&lt;/p&gt;</description></item><item><title>About</title><link>https://bertobi.github.io/blog/about/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://bertobi.github.io/blog/about/</guid><description>&lt;p&gt;Hi, I&amp;rsquo;m &lt;strong&gt;Tobias&lt;/strong&gt; — a junior AI Safety researcher based in Buenos Aires.&lt;/p&gt;
&lt;p&gt;I started this blog to share my work in AI Safety and ocasional showcases of personal projects.&lt;/p&gt;
&lt;p&gt;You can find me on &lt;a href="https://github.com/BerTobi"&gt;GitHub&lt;/a&gt;. Say hello in the comments
on any post.&lt;/p&gt;</description></item></channel></rss>