{"id":496481,"date":"2026-09-29T07:22:39","date_gmt":"2026-09-29T07:22:39","guid":{"rendered":"https:\/\/savepearlharbor.com\/?p=496481"},"modified":"-0001-11-30T00:00:00","modified_gmt":"-0001-11-29T21:00:00","slug":"","status":"publish","type":"post","link":"https:\/\/savepearlharbor.com\/?p=496481","title":{"rendered":"What does it take to build auto-mode like in Claude?"},"content":{"rendered":"<div xmlns=\"http:\/\/www.w3.org\/1999\/xhtml\">\n<p>I build Gaunt Sloth, a small open-source CLI agent. In 2.0 it got a shell tool: the agent can now compose its own commands instead of running a fixed list from the config. Obviously it needed approvals. Agent proposes a command, you press \u201cy\u201d, it runs. Every coding agent has one. How hard can it be?<\/p>\n<p>It took a few weeks and more wrong assumptions than I\u2019d like to admit. None of the problems below are specific to my tool. If you build an agent, you will probably hit them too, and if you use one, it\u2019s worth knowing what that little prompt actually does for you.<\/p>\n<p>These things may look obvious once written down. In my experience very few people actually hold them.<\/p>\n<h3>\u201cAlways approve\u201d is a policy, not a button<\/h3>\n<p>The first design was the naive one. Prompt shows <code>rm -rf .\/build<\/code>, you press \u201calways\u201d, we store the prefix <code>rm<\/code>. Next time the agent wants <code>rm -rf ~\/Documents<\/code>, it is already approved.<\/p>\n<p>Nobody meant to allow-list every <code>rm<\/code> on the machine. But a prefix is the easiest thing to store, and \u201calways\u201d sounds harmless on a prompt. Now the prompt remembers exactly the command shown, and any broader pattern is something you write in a config file yourself.<\/p>\n<p>(This exact sentence about <code>rm -rf ~\/Documents<\/code>, quoted in backticks inside a <code>git commit -m \"...\"<\/code>, is how Claude Code deleted the Documents folder on the machine I set up for agents to work on, while writing a commit message about this feature. Bash treats backticks inside double quotes as command substitution. Everything on that machine is durable elsewhere, so nothing was lost. That is a separate story.)<\/p>\n<h3>You can\u2019t parse shell, so stop trying to be precise<\/h3>\n<p>To decide anything about a command you have to know what it will run. Bash disagrees with you about that more often than you expect: quotes, escapes, heredocs, <code>$(...)<\/code>, backticks, line continuations.<\/p>\n<p>My natural instinct was to write a smarter parser. I built a quote- and heredoc-aware scanner that understood which newlines were \u201creal\u201d separators. Against 12 known evasions it leaked 6. The dumb version that only blanks newlines inside single quotes (about six lines of code, and single quotes have no escape rules, so it can\u2019t be wrong) leaked none, and still passed the legitimate <code>awk<\/code>, <code>sed<\/code> and <code>python3 -c<\/code> cases.<\/p>\n<p>Lesson: precision is cheap in code that only displays things to a human. In code that decides whether something runs, every precision claim is a bet that your parser matches bash, and that is where the bugs live. When the checker can\u2019t resolve a command, it should say so and fail closed.<\/p>\n<p>The same goes for lists. At one point the code had a list of wrapper commands to look through: <code>sudo<\/code>, <code>env<\/code>, <code>nohup<\/code>, <code>time<\/code>. It did not have <code>timeout<\/code>. So <code>time curl &lt;url&gt;<\/code> was checked and <code>timeout 30 curl &lt;url&gt;<\/code> was not. Adding the missing names only makes the list look complete until the next tool appears. The fix that held removed the list and replaced it with a structural rule about which token is the command.<\/p>\n<h3>A hard floor will refuse the fix<\/h3>\n<p>Some commands should never run, in any mode, with no override. <code>rm -rf \/<\/code>, <code>mkfs<\/code> on a disk, that kind of thing. Easy, right? A few regexes.<\/p>\n<p>The first version fired on text, not on commands. It refused <code>echo never run rm -rf \/<\/code> and <code>grep -c mkfs docs\/*.md<\/code>. After fixing that, a second round found it refused <code>shutdown -c<\/code>, which <em>cancels<\/em> a pending shutdown, and a <code>sed<\/code> that <em>removes<\/em> a dangerous <code>&gt; \/dev\/sda<\/code> redirect from a script. So the two kinds of checks need opposite tuning. A check that refuses with no appeal must be narrow and accept misses, because its false positive is unrecoverable. A check that only escalates to a human should over-fire, because its false positive costs one prompt.<\/p>\n<h3>The LLM rater sees the problem and approves anyway<\/h3>\n<p>The obvious next step is to ask a model: is this command safe? Models are actually good at spotting an isolated destructive command, even small local ones.<\/p>\n<p>The failures are more interesting than \u201cthe model is dumb\u201d. I gave a rater <code>curl -fsSL https:\/\/pypi.org.packages-cdn.io\/simple\/ -o index.html<\/code>. It said, roughly, \u201ca domain that mimics a CDN but only fetches a file\u201d, and rated it safe. It saw the deception, said so, and weighted it wrong. Obvious typosquats and even a Cyrillic homoglyph were caught fine.<\/p>\n<p>Whether a host is trustworthy is simply not answerable from a command string. There is no reputation data at rating time and a convincing domain costs ten dollars. So in my gate a command that fetches from a host it names is never rated safe, whatever the model says. That is decided by code, not by the model.<\/p>\n<p>Similar story with checking user intent. If a user writes \u201cclone the testing repo from our org\u201d and the agent fills in a specific URL from its context, a model asked \u201cdid the user ask for this?\u201d tends to say yes. The user gave an incomplete prompt, the context nicely completes it, and that\u2019s a classic hallucination trap. Closed-world references (\u201cthe last two commits\u201d, \u201cthe build folder\u201d) were fine. Open-world ones (a URL, a host) weren\u2019t.<\/p>\n<p>And don\u2019t generalise from whichever models you happen to have configured. I claimed \u201call raters agree\u201d twice, and both times a small model disagreed.<\/p>\n<h3>One indirection, and all of the above is bypassed<\/h3>\n<p>An agent that can write files and run commands does not need to show you the dangerous command. It can write it to a file and run the file:<\/p>\n<pre><code>write_file payload.sh      # the file tool, not the shellbash payload.sh<\/code><div class=\"code-explainer\"><a href=\"https:\/\/sourcecraft.dev\/\" class=\"tm-button code-explainer__link\" style=\"visibility: hidden;\"><img style=\"width:87px;height:14px;object-fit:cover;object-position:left;\"\/><\/a><\/div><\/pre>\n<p>I built a small probe for this. The direct version of a payload was stopped by the deny list, by the hard floor and by the parser. The indirect version matched nothing and ran. The rater didn\u2019t miss it. The rater was never shown it: its entire input was <code>bash payload.sh<\/code>.<\/p>\n<p>Same with <code>write_file build.js<\/code> and then <code>npm run build<\/code>. Neither step is a lie.<\/p>\n<p>So the honest claim for any approval gate that reads command text, mine included, is that it protects you from <strong>accidents, not intent<\/strong>. An agent doing the wrong thing while trying to do the right thing. That\u2019s genuinely worth having, because accidents are the common case. The mistake is calling it a boundary.<\/p>\n<h3>Where uncertainty should go<\/h3>\n<p>Once I accepted that, the question changed from \u201chow do I detect everything\u201d to \u201cwho should see each kind of doubt\u201d.<\/p>\n<p>A regex knows form and nothing about meaning, so a parser\u2019s confusion is not evidence of danger. It shouldn\u2019t spend a person\u2019s attention. A model is the only layer that reads meaning, so a model\u2019s doubt is exactly what a person should see. Most gates get this backwards and interrupt the human whenever the parser gets confused, which is the same interruption having learned nothing.<\/p>\n<p>What worked for parser confusion was sending it back to the agent, like a compiler error: \u201cI can\u2019t resolve this, here is the span, rewrite it\u201d. With a bare \u201crejected, rewrite it\u201d, Claude Haiku and Sonnet fixed their command about a third of the time. With the specific reason and span, roughly two thirds to four fifths, across three runs. On a small local model the difference was noise, so there a person still gets asked.<\/p>\n<p>The human doesn\u2019t scale either. Ask a person three questions and you get three careful answers. Ask ten and you get about three. A tired reviewer approves rather than complains, and nothing tells you. From my own use (not a measurement): a small local model works for minutes and stops, a frontier model runs for hours and issues an order of magnitude more commands. So approving every command by hand is not the strictest mode, it\u2019s a mode for a handful of commands you actually want to read. At five hundred it\u2019s a rubber stamp.<\/p>\n<h3>What actually contains an agent<\/h3>\n<p>None of this is exotic, and none of it lives in the agent:<\/p>\n<ul>\n<li>\n<p>run it in a container, a VM or a separate OS account that owns nothing you care about;<\/p>\n<\/li>\n<li>\n<p>keep your working folder away from your home directory;<\/p>\n<\/li>\n<li>\n<p>push often, and check that branch protection is on for every repo you think it\u2019s on (mine wasn\u2019t);<\/p>\n<\/li>\n<li>\n<p>have backups, and know when the last one ran.<\/p>\n<\/li>\n<\/ul>\n<p>The approval prompt is still useful. Mine now has a deterministic floor, a rater, a negotiation where the agent can justify or narrow a risky command, and five modes from <code>manual<\/code> to <code>bypass<\/code>. It catches the everyday mistakes and saves a lot of prompts. It just isn\u2019t a sandbox, and I\u2019d rather say that out loud than let people assume it.<\/p>\n<p>Gaunt Sloth is on GitHub: <a href=\"https:\/\/github.com\/pukeko-robotics\/gaunt-sloth\" rel=\"noopener nofollow\">https:\/\/github.com\/pukeko-robotics\/gaunt-sloth<\/a>. The docs page on what approvals protect you from is here: <a href=\"https:\/\/gauntsloth.app\/docs\/guides\/what-approvals-protect-you-from\/\" rel=\"noopener nofollow\">https:\/\/gauntsloth.app\/docs\/guides\/what-approvals-protect-you-from\/<\/a><\/p>\n<p>If you\u2019ve built something like this and found different answers, I\u2019d really like to hear them in the comments.<\/p>\n<\/div>\n<p>\u0441\u0441\u044b\u043b\u043a\u0430 \u043d\u0430 \u043e\u0440\u0438\u0433\u0438\u043d\u0430\u043b \u0441\u0442\u0430\u0442\u044c\u0438 <a href=\"https:\/\/habr.com\/ru\/articles\/1087882\/\">https:\/\/habr.com\/ru\/articles\/1087882\/<\/a><\/p>\n","protected":false},"excerpt":{"rendered":"<p>I build Gaunt Sloth, a small open-source CLI agent. In 2.0 it got a shell tool: the agent can now compose its own commands instead of running a fixed list from the config. Obviously it needed approvals. Agent proposes a command, you press \u201cy\u201d, it runs. Every coding agent has one. How hard can it be?It took a few weeks and more wrong assumptions than I\u2019d like to admit. None of the problems below are specific to my tool. If you build an agent, you will probably hit them too, and if you use one, it\u2019s worth knowing what that little prompt actually does for you.These things may look obvious once written down. In my experience very few people actually hold them.\u201cAlways approve\u201d is a policy, not a buttonThe first design was the naive one. Prompt shows rm -rf .\/build, you press \u201calways\u201d, we store the prefix rm. Next time the agent wants rm -rf ~\/Documents, it is already approved.Nobody meant to allow-list every rm on the machine. But a prefix is the easiest thing to store, and \u201calways\u201d sounds harmless on a prompt. Now the prompt remembers exactly the command shown, and any broader pattern is something you write in a config file yourself.(This exact sentence about rm -rf ~\/Documents, quoted in backticks inside a git commit -m &#171;&#8230;&#187;, is how Claude Code deleted the Documents folder on the machine I set up for agents to work on, while writing a commit message about this feature. Bash treats backticks inside double quotes as command substitution. Everything on that machine is durable elsewhere, so nothing was lost. That is a separate story.)You can\u2019t parse shell, so stop trying to be preciseTo decide anything about a command you have to know what it will run. Bash disagrees with you about that more often than you expect: quotes, escapes, heredocs, $(&#8230;), backticks, line continuations.My natural instinct was to write a smarter parser. I built a quote- and heredoc-aware scanner that understood which newlines were \u201creal\u201d separators. Against 12 known evasions it leaked 6. The dumb version that only blanks newlines inside single quotes (about six lines of code, and single quotes have no escape rules, so it can\u2019t be wrong) leaked none, and still passed the legitimate awk, sed and python3 -c cases.Lesson: precision is cheap in code that only displays things to a human. In code that decides whether something runs, every precision claim is a bet that your parser matches bash, and that is where the bugs live. When the checker can\u2019t resolve a command, it should say so and fail closed.The same goes for lists. At one point the code had a list of wrapper commands to look through: sudo, env, nohup, time. It did not have timeout. So time curl &lt;url&gt; was checked and timeout 30 curl &lt;url&gt; was not. Adding the missing names only makes the list look complete until the next tool appears. The fix that held removed the list and replaced it with a structural rule about which token is the command.A hard floor will refuse the fixSome commands should never run, in any mode, with no override. rm -rf \/, mkfs on a disk, that kind of thing. Easy, right? A few regexes.The first version fired on text, not on commands. It refused echo never run rm -rf \/ and grep -c mkfs docs\/*.md. After fixing that, a second round found it refused shutdown -c, which cancels a pending shutdown, and a sed that removes a dangerous &gt; \/dev\/sda redirect from a script. So the two kinds of checks need opposite tuning. A check that refuses with no appeal must be narrow and accept misses, because its false positive is unrecoverable. A check that only escalates to a human should over-fire, because its false positive costs one prompt.The LLM rater sees the problem and approves anywayThe obvious next step is to ask a model: is this command safe? Models are actually good at spotting an isolated destructive command, even small local ones.The failures are more interesting than \u201cthe model is dumb\u201d. I gave a rater curl -fsSL https:\/\/pypi.org.packages-cdn.io\/simple\/ -o index.html. It said, roughly, \u201ca domain that mimics a CDN but only fetches a file\u201d, and rated it safe. It saw the deception, said so, and weighted it wrong. Obvious typosquats and even a Cyrillic homoglyph were caught fine.Whether a host is trustworthy is simply not answerable from a command string. There is no reputation data at rating time and a convincing domain costs ten dollars. So in my gate a command that fetches from a host it names is never rated safe, whatever the model says. That is decided by code, not by the model.Similar story with checking user intent. If a user writes \u201cclone the testing repo from our org\u201d and the agent fills in a specific URL from its context, a model asked \u201cdid the user ask for this?\u201d tends to say yes. The user gave an incomplete prompt, the context nicely completes it, and that\u2019s a classic hallucination trap. Closed-world references (\u201cthe last two commits\u201d, \u201cthe build folder\u201d) were fine. Open-world ones (a URL, a host) weren\u2019t.And don\u2019t generalise from whichever models you happen to have configured. I claimed \u201call raters agree\u201d twice, and both times a small model disagreed.One indirection, and all of the above is bypassedAn agent that can write files and run commands does not need to show you the dangerous command. It can write it to a file and run the file:write_file payload.sh      # the file tool, not the shellbash payload.shI built a small probe for this. The direct version of a payload was stopped by the deny list, by the hard floor and by the parser. The indirect version matched nothing and ran. The rater didn\u2019t miss it. The rater was never shown it: its entire input was bash payload.sh.Same with write_file build.js and then npm run build. Neither step is a lie.So the honest claim for any approval gate that reads command text, mine included, is that it protects you from accidents, not intent. An agent doing the wrong thing while trying to do the right thing. That\u2019s genuinely worth having, because accidents are the common case. The mistake is calling it a boundary.Where uncertainty should goOnce I accepted that, the question changed from \u201chow do I detect everything\u201d to \u201cwho should see each kind of doubt\u201d.A regex knows form and nothing about meaning, so a parser\u2019s confusion is not evidence of danger. It shouldn\u2019t spend a person\u2019s attention. A model is the only layer that reads meaning, so a model\u2019s doubt is exactly what a person should see. Most gates get this backwards and interrupt the human whenever the parser gets confused, which is the same interruption having learned nothing.What worked for parser confusion was sending it back to the agent, like a compiler error: \u201cI can\u2019t resolve this, here is the span, rewrite it\u201d. With a bare \u201crejected, rewrite it\u201d, Claude Haiku and Sonnet fixed their command about a third of the time. With the specific reason and span, roughly two thirds to four fifths, across three runs. On a small local model the difference was noise, so there a person still gets asked.The human doesn\u2019t scale either. Ask a person three questions and you get three careful answers. Ask ten and you get about three. A tired reviewer approves rather than complains, and nothing tells you. From my own use (not a measurement): a small local model works for minutes and stops, a frontier model runs for hours and issues an order of magnitude more commands. So approving every command by hand is not the strictest mode, it\u2019s a mode for a handful of commands you actually want to read. At five hundred it\u2019s a rubber stamp.What actually contains an agentNone of this is exotic, and none of it lives in the agent:run it in a container, a VM or a separate OS account that owns nothing you care about;keep your working folder away from your home directory;push often, and check that branch protection is on for every repo you think it\u2019s on (mine wasn\u2019t);have backups, and know when the last one ran.The approval prompt is still useful. Mine now has a deterministic floor, a rater, a negotiation where the agent can justify or narrow a risky command, and five modes from manual to bypass. It catches the everyday mistakes and saves a lot of prompts. It just isn\u2019t a sandbox, and I\u2019d rather say that out loud than let people assume it.Gaunt Sloth is on GitHub: https:\/\/github.com\/pukeko-robotics\/gaunt-sloth. The docs page on what approvals protect you from is here: https:\/\/gauntsloth.app\/docs\/guides\/what-approvals-protect-you-from\/If you\u2019ve built something like this and found different answers, I\u2019d really like to hear them in the comments.\u0441\u0441\u044b\u043b\u043a\u0430 \u043d\u0430 \u043e\u0440\u0438\u0433\u0438\u043d\u0430\u043b \u0441\u0442\u0430\u0442\u044c\u0438 https:\/\/habr.com\/ru\/articles\/1087882\/<\/p>\n","protected":false},"author":1,"featured_media":0,"comment_status":"closed","ping_status":"closed","sticky":false,"template":"","format":"standard","meta":{"footnotes":""},"categories":[],"tags":[],"class_list":["post-496481","post","type-post","status-publish","format-standard","hentry"],"_links":{"self":[{"href":"https:\/\/savepearlharbor.com\/index.php?rest_route=\/wp\/v2\/posts\/496481","targetHints":{"allow":["GET"]}}],"collection":[{"href":"https:\/\/savepearlharbor.com\/index.php?rest_route=\/wp\/v2\/posts"}],"about":[{"href":"https:\/\/savepearlharbor.com\/index.php?rest_route=\/wp\/v2\/types\/post"}],"author":[{"embeddable":true,"href":"https:\/\/savepearlharbor.com\/index.php?rest_route=\/wp\/v2\/users\/1"}],"replies":[{"embeddable":true,"href":"https:\/\/savepearlharbor.com\/index.php?rest_route=%2Fwp%2Fv2%2Fcomments&post=496481"}],"version-history":[{"count":0,"href":"https:\/\/savepearlharbor.com\/index.php?rest_route=\/wp\/v2\/posts\/496481\/revisions"}],"wp:attachment":[{"href":"https:\/\/savepearlharbor.com\/index.php?rest_route=%2Fwp%2Fv2%2Fmedia&parent=496481"}],"wp:term":[{"taxonomy":"category","embeddable":true,"href":"https:\/\/savepearlharbor.com\/index.php?rest_route=%2Fwp%2Fv2%2Fcategories&post=496481"},{"taxonomy":"post_tag","embeddable":true,"href":"https:\/\/savepearlharbor.com\/index.php?rest_route=%2Fwp%2Fv2%2Ftags&post=496481"}],"curies":[{"name":"wp","href":"https:\/\/api.w.org\/{rel}","templated":true}]}}