<?xml version="1.0" encoding="utf-8" standalone="yes"?><rss version="2.0" xmlns:atom="http://www.w3.org/2005/Atom"><channel><title>Thoughts on Sina's Page</title><link>https://sina.page/thoughts/</link><description>Recent content in Thoughts on Sina's Page</description><generator>Hugo</generator><language>en</language><copyright/><lastBuildDate>Sun, 09 Aug 2026 00:00:00 +0000</lastBuildDate><atom:link href="https://sina.page/thoughts/index.xml" rel="self" type="application/rss+xml"/><item><title>The New Delegation Stack</title><link>https://sina.page/thoughts/delegation-stack/</link><pubDate>Sun, 09 Aug 2026 00:00:00 +0000</pubDate><guid>https://sina.page/thoughts/delegation-stack/</guid><description>&lt;p&gt;Somewhere in Anthropic&amp;rsquo;s production stack there is a prompt that teaches one of the most capable models on Earth how to be a middle manager. It spells out staffing policy by hand: simple fact-finding gets one agent and three to ten tool calls; a direct comparison gets two to four subagents with ten to fifteen calls each; complex research gets ten or more subagents with clearly divided responsibilities (&lt;a href="https://www.anthropic.com/engineering/multi-agent-research-system" target="_blank" rel="noreferrer"&gt;How we built our multi-agent research system&lt;/a&gt;). They wrote those rules because, without them, the lead agent did what every new manager does — early versions spawned fifty subagents for simple queries (the review&amp;rsquo;s account of the system).&lt;/p&gt;</description></item><item><title>Evals Are the New Engineering Artifacts</title><link>https://sina.page/thoughts/evals-engineering-artifacts/</link><pubDate>Sun, 31 May 2026 00:00:00 +0000</pubDate><guid>https://sina.page/thoughts/evals-engineering-artifacts/</guid><description>&lt;p&gt;Traditional software relies on explicit specifications. AI systems increasingly require a different contract: the ability to measure whether behavior is good.&lt;/p&gt;
&lt;p&gt;Evaluation becomes a core engineering artifact.&lt;/p&gt;
&lt;p&gt;It defines what success means, creates feedback loops, and allows intelligent systems to improve.&lt;/p&gt;</description></item></channel></rss>