<?xml version="1.0" encoding="utf-8" standalone="yes"?>
<rss version="2.0" xmlns:atom="http://www.w3.org/2005/Atom">
  <channel>
    <title>Reliability on aeshift</title>
    <link>https://aeshift.com/tags/reliability/</link>
    <description>Recent content in Reliability on aeshift</description>
    <generator>Hugo -- gohugo.io</generator>
    <language>en</language>
    <copyright>&amp;copy; 2026 [Dachary Carey](https://dacharycarey.com) - with agent assistance · Part of the [Agent Ecosystem Research Program](https://agentecosystem.dev)</copyright>
    <lastBuildDate>Wed, 18 Mar 2026 00:00:00 +0000</lastBuildDate><atom:link href="https://aeshift.com/tags/reliability/index.xml" rel="self" type="application/rss+xml" />
    
    <item>
      <title>AI Agents Have Stable &#39;Coding Styles&#39; That Change With Each Version</title>
      <link>https://aeshift.com/posts/2026-03-18-nonstandard-errors-in-ai-agents/</link>
      <pubDate>Wed, 18 Mar 2026 00:00:00 +0000</pubDate>
      
      <guid>https://aeshift.com/posts/2026-03-18-nonstandard-errors-in-ai-agents/</guid>
      <description>&lt;p&gt;If you&amp;rsquo;re using coding agents to produce analysis, you&amp;rsquo;re not running deterministic software. You&amp;rsquo;re managing a lab: multiple researchers with consistent &amp;ldquo;styles,&amp;rdquo; inconsistent choices, and outcomes that drift even when the prompt and data don&amp;rsquo;t.&lt;/p&gt;&#xA;&lt;p&gt;The authors of &lt;em&gt;&lt;a href=&#34;http://arxiv.org/abs/2603.16744v1&#34;  target=&#34;_blank&#34; rel=&#34;noreferrer&#34;&gt;Nonstandard Errors in AI Agents&lt;/a&gt;&lt;/em&gt; ran 150 autonomous Claude Code agents on the same NYSE TAQ dataset (SPY, 2015–2024) and the same six hypotheses. The results varied because the agents made different methodological choices, and those choices often &lt;em&gt;are&lt;/em&gt; the analysis.&lt;/p&gt;</description>
      <media:content xmlns:media="http://search.yahoo.com/mrss/" url="https://aeshift.com/posts/2026-03-18-nonstandard-errors-in-ai-agents/feature.png" />
    </item>
    
    <item>
      <title>Don&#39;t Let Your Agent Grade Its Own Homework</title>
      <link>https://aeshift.com/posts/2026-03-06-self-attribution-bias-when-ai-monitors-go-easy-on-themselves/</link>
      <pubDate>Fri, 06 Mar 2026 00:00:00 +0000</pubDate>
      
      <guid>https://aeshift.com/posts/2026-03-06-self-attribution-bias-when-ai-monitors-go-easy-on-themselves/</guid>
      <description>&lt;p&gt;If you&amp;rsquo;re using an LLM to monitor an LLM-based coding agent, assume the monitor is biased in favor of the agent&amp;rsquo;s own output. The evidence suggests that framing matters: the same risky action looks safer when it&amp;rsquo;s presented as something the assistant just did.&lt;/p&gt;&#xA;&lt;p&gt;That&amp;rsquo;s the core finding of &lt;a href=&#34;http://arxiv.org/abs/2603.04582v1&#34;  target=&#34;_blank&#34; rel=&#34;noreferrer&#34;&gt;&amp;ldquo;Self-Attribution Bias: When AI Monitors Go Easy on Themselves&amp;rdquo;&lt;/a&gt;. For practitioners, this is less an AI psychology curiosity and more an engineering warning: self-monitoring setups can systematically under-flag the exact failures you&amp;rsquo;re trying to catch.&lt;/p&gt;</description>
      <media:content xmlns:media="http://search.yahoo.com/mrss/" url="https://aeshift.com/posts/2026-03-06-self-attribution-bias-when-ai-monitors-go-easy-on-themselves/feature.jpg" />
    </item>
    
  </channel>
</rss>
