<?xml version="1.0" encoding="utf-8" standalone="yes"?><rss version="2.0" xmlns:atom="http://www.w3.org/2005/Atom" xmlns:content="http://purl.org/rss/1.0/modules/content/"><channel><title>Blog — Ankush Kun</title><link>https://ankush.one/blogs/</link><description>Ankush's blog - Thoughts, tutorials, and learnings about web development, APIs, CI/CD, and open-source software.</description><generator>Hugo -- gohugo.io</generator><language>en</language><managingEditor>ankush4singh@gmail.com (Ankush Singh)</managingEditor><webMaster>ankush4singh@gmail.com (Ankush Singh)</webMaster><lastBuildDate>Mon, 12 Oct 2026 02:59:21 +0530</lastBuildDate><atom:link href="https://ankush.one/blogs/index.xml" rel="self" type="application/rss+xml"/><item><title>How to switch Claude Code accounts without losing your conversation</title><link>https://ankush.one/blogs/oof-claude-code-account-switcher/</link><pubDate>Mon, 12 Oct 2026 01:00:00 +0530</pubDate><author>ankush4singh@gmail.com (Ankush Singh)</author><guid>https://ankush.one/blogs/oof-claude-code-account-switcher/</guid><category>claude-code</category><category>developer-tools</category><category>ai-agents</category><description>I kept losing my flow to the /login browser dance whenever I needed a different Claude Code account mid-task. So I wrote oof, a small zsh script that saves the current login, swaps in another one and resumes the same conversation.</description><content:encoded>&lt;p&gt;Links: &lt;a href="https://gist.github.com/ankushKun/cd27fb07b0f06be99d4181b40309103d"&gt;oof on GitHub Gist&lt;/a&gt;&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Short answer:&lt;/strong&gt; Claude Code keeps your conversations on disk per folder, not per account. So you can swap the login, run &lt;code&gt;claude --continue&lt;/code&gt; in the same folder, and keep going with the same conversation on a different account. &lt;code&gt;oof&lt;/code&gt; does that in one command. It saves the current login from the macOS Keychain, swaps in the account you asked for, updates &lt;code&gt;~/.claude.json&lt;/code&gt; and resumes. The whole script is &lt;a href="https://gist.github.com/ankushKun/cd27fb07b0f06be99d4181b40309103d" data-external&gt;on GitHub Gist&lt;/a&gt;.&lt;/p&gt;
&lt;p&gt;I use Claude Code for &lt;a href="https://construct.computer" data-external&gt;Construct&lt;/a&gt;, my startup, and for side projects like this site, and I have more than one Claude account. When I&amp;rsquo;m deep in a session and need a different one, the built-in way is &lt;code&gt;/login&lt;/code&gt;: a browser tab opens, I check which account the browser is signed in to, approve, come back to the terminal and hope I picked the right one. It works, but it gets old fast, and it always seems to happen right when I&amp;rsquo;m in the middle of something.&lt;/p&gt;
&lt;p&gt;What I wanted was one command that says &amp;ldquo;use the other account&amp;rdquo; and puts me back in the exact conversation I was having. No browser, no copying tokens around, and no fresh chat where I have to explain the whole task again.&lt;/p&gt;
&lt;p&gt;That command is &lt;code&gt;oof&lt;/code&gt;, named after the noise I make when Claude Code tells me I&amp;rsquo;ve hit a limit halfway through a task. It&amp;rsquo;s a single zsh script that I wrote in one sitting. This post covers how it works, how it compares with the official way of running several accounts, and the small rabbit holes I fell into along the way.&lt;/p&gt;
&lt;h2 id="where-does-claude-code-store-your-login"&gt;Where does Claude Code store your login?&lt;a class="heading-anchor" href="#where-does-claude-code-store-your-login" aria-label="Link to this section"&gt;&lt;/a&gt;
&lt;/h2&gt;
&lt;p&gt;The key fact is that &lt;strong&gt;your conversations don&amp;rsquo;t belong to your account&lt;/strong&gt;. Claude Code keeps session history on disk under &lt;code&gt;~/.claude/projects/&lt;/code&gt;, grouped by the folder you ran it in. The account only decides who pays for the next message. If you swap the login and run &lt;code&gt;claude -c&lt;/code&gt; in the same folder, you land back in your most recent conversation, just on a different account.&lt;/p&gt;
&lt;p&gt;That leaves the question of where the login itself lives. On macOS it&amp;rsquo;s split across two places:&lt;/p&gt;
&lt;div class="highlight"&gt;&lt;pre tabindex="0" class="chroma"&gt;&lt;code class="language-text" data-lang="text"&gt;&lt;span class="line"&gt;&lt;span class="cl"&gt; macOS Keychain ~/.claude.json
&lt;/span&gt;&lt;/span&gt;&lt;span class="line"&gt;&lt;span class="cl"&gt; ┌──────────────────────────────┐ ┌──────────────────────────────┐
&lt;/span&gt;&lt;/span&gt;&lt;span class="line"&gt;&lt;span class="cl"&gt; │ &amp;#34;Claude Code-credentials&amp;#34; │ │ &amp;#34;oauthAccount&amp;#34;: { │
&lt;/span&gt;&lt;/span&gt;&lt;span class="line"&gt;&lt;span class="cl"&gt; │ OAuth access + refresh │ │ &amp;#34;emailAddress&amp;#34;: ..., │
&lt;/span&gt;&lt;/span&gt;&lt;span class="line"&gt;&lt;span class="cl"&gt; │ tokens, plan, rate tier │ │ &amp;#34;organizationUuid&amp;#34;: ... │
&lt;/span&gt;&lt;/span&gt;&lt;span class="line"&gt;&lt;span class="cl"&gt; │ (the secret) │ │ } (who /status says you are)│
&lt;/span&gt;&lt;/span&gt;&lt;span class="line"&gt;&lt;span class="cl"&gt; └──────────────────────────────┘ └──────────────────────────────┘
&lt;/span&gt;&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;&lt;/div&gt;&lt;p&gt;The Keychain item holds the actual OAuth tokens. The &lt;code&gt;oauthAccount&lt;/code&gt; block in &lt;code&gt;~/.claude.json&lt;/code&gt; is just the label, the email and organization that &lt;code&gt;/status&lt;/code&gt; shows. To switch accounts you have to change both, and to switch back later you need a copy of each account&amp;rsquo;s tokens stored somewhere safe. On Linux and Windows the tokens live in &lt;code&gt;~/.claude/.credentials.json&lt;/code&gt; instead, so the same idea becomes a file copy.&lt;/p&gt;
&lt;h2 id="why-not-login-or-claude_config_dir"&gt;Why not /login or CLAUDE_CONFIG_DIR?&lt;a class="heading-anchor" href="#why-not-login-or-claude_config_dir" aria-label="Link to this section"&gt;&lt;/a&gt;
&lt;/h2&gt;
&lt;p&gt;There are already two ways to use multiple Claude Code accounts, and it&amp;rsquo;s worth being clear about why I didn&amp;rsquo;t stop at either.&lt;/p&gt;
&lt;p&gt;&lt;code&gt;/login&lt;/code&gt; works inside a running session and keeps the conversation, but it sends you through the browser every single time.&lt;/p&gt;
&lt;p&gt;The official answer for multiple accounts is &lt;code&gt;CLAUDE_CONFIG_DIR&lt;/code&gt;. When someone asked for named account profiles in &lt;a href="https://github.com/anthropics/claude-code/issues/27359" data-external&gt;anthropics/claude-code#27359&lt;/a&gt;, a maintainer closed it by pointing out that a separate config directory per account already gives you this, with an alias for each:&lt;/p&gt;
&lt;div class="highlight"&gt;&lt;pre tabindex="0" class="chroma"&gt;&lt;code class="language-zsh" data-lang="zsh"&gt;&lt;span class="line"&gt;&lt;span class="cl"&gt;&lt;span class="nb"&gt;alias&lt;/span&gt; claude-work&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="s1"&gt;&amp;#39;CLAUDE_CONFIG_DIR=~/.claude-work claude&amp;#39;&lt;/span&gt;
&lt;/span&gt;&lt;/span&gt;&lt;span class="line"&gt;&lt;span class="cl"&gt;&lt;span class="nb"&gt;alias&lt;/span&gt; claude-personal&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="s1"&gt;&amp;#39;CLAUDE_CONFIG_DIR=~/.claude-personal claude&amp;#39;&lt;/span&gt;
&lt;/span&gt;&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;&lt;/div&gt;&lt;p&gt;Each directory has its own login, settings and &lt;code&gt;projects/&lt;/code&gt; folder, which is exactly the problem for me. My conversation lives in one directory, so when I move to the other account it isn&amp;rsquo;t there to continue. You can symlink the history folders together, but then you&amp;rsquo;re maintaining a setup that fights the tool&amp;rsquo;s own design.&lt;/p&gt;
&lt;table&gt;
	&lt;thead&gt;
			&lt;tr&gt;
					&lt;th&gt;&lt;/th&gt;
					&lt;th&gt;&lt;code&gt;/login&lt;/code&gt;&lt;/th&gt;
					&lt;th&gt;&lt;code&gt;CLAUDE_CONFIG_DIR&lt;/code&gt; per account&lt;/th&gt;
					&lt;th&gt;&lt;code&gt;oof&lt;/code&gt;&lt;/th&gt;
			&lt;/tr&gt;
	&lt;/thead&gt;
	&lt;tbody&gt;
			&lt;tr&gt;
					&lt;td&gt;Browser sign-in on every switch&lt;/td&gt;
					&lt;td&gt;Yes&lt;/td&gt;
					&lt;td&gt;No, once per directory&lt;/td&gt;
					&lt;td&gt;No, once per account&lt;/td&gt;
			&lt;/tr&gt;
			&lt;tr&gt;
					&lt;td&gt;Continue the same conversation&lt;/td&gt;
					&lt;td&gt;Yes&lt;/td&gt;
					&lt;td&gt;No, each directory has its own history&lt;/td&gt;
					&lt;td&gt;Yes, through &lt;code&gt;claude -c&lt;/code&gt;&lt;/td&gt;
			&lt;/tr&gt;
			&lt;tr&gt;
					&lt;td&gt;Settings, MCP servers and memory&lt;/td&gt;
					&lt;td&gt;Shared&lt;/td&gt;
					&lt;td&gt;Separate per account&lt;/td&gt;
					&lt;td&gt;Shared&lt;/td&gt;
			&lt;/tr&gt;
			&lt;tr&gt;
					&lt;td&gt;Two accounts running at once&lt;/td&gt;
					&lt;td&gt;No&lt;/td&gt;
					&lt;td&gt;Yes&lt;/td&gt;
					&lt;td&gt;No, one login is active&lt;/td&gt;
			&lt;/tr&gt;
			&lt;tr&gt;
					&lt;td&gt;Setup&lt;/td&gt;
					&lt;td&gt;None&lt;/td&gt;
					&lt;td&gt;An alias per account&lt;/td&gt;
					&lt;td&gt;One script and two hooks&lt;/td&gt;
			&lt;/tr&gt;
	&lt;/tbody&gt;
&lt;/table&gt;
&lt;p&gt;If you want strict separation, for example two clients whose work should never mix, or you need two accounts live at the same time, use &lt;code&gt;CLAUDE_CONFIG_DIR&lt;/code&gt;. If you want one setup and the freedom to carry the same piece of work across accounts, that&amp;rsquo;s what &lt;code&gt;oof&lt;/code&gt; is for.&lt;/p&gt;
&lt;h2 id="version-one-two-shell-functions"&gt;Version one: two shell functions&lt;a class="heading-anchor" href="#version-one-two-shell-functions" aria-label="Link to this section"&gt;&lt;/a&gt;
&lt;/h2&gt;
&lt;p&gt;The first idea was two functions in &lt;code&gt;.zshrc&lt;/code&gt;:&lt;/p&gt;
&lt;div class="highlight"&gt;&lt;pre tabindex="0" class="chroma"&gt;&lt;code class="language-zsh" data-lang="zsh"&gt;&lt;span class="line"&gt;&lt;span class="cl"&gt;cc-save&lt;span class="o"&gt;()&lt;/span&gt; &lt;span class="o"&gt;{&lt;/span&gt; &lt;span class="c1"&gt;# cc-save work&lt;/span&gt;
&lt;/span&gt;&lt;/span&gt;&lt;span class="line"&gt;&lt;span class="cl"&gt; security add-generic-password -U -a &lt;span class="s2"&gt;&amp;#34;&lt;/span&gt;&lt;span class="nv"&gt;$USER&lt;/span&gt;&lt;span class="s2"&gt;&amp;#34;&lt;/span&gt; -s &lt;span class="s2"&gt;&amp;#34;Claude Code-credentials-&lt;/span&gt;&lt;span class="nv"&gt;$1&lt;/span&gt;&lt;span class="s2"&gt;&amp;#34;&lt;/span&gt; &lt;span class="se"&gt;\
&lt;/span&gt;&lt;/span&gt;&lt;/span&gt;&lt;span class="line"&gt;&lt;span class="cl"&gt; -w &lt;span class="s2"&gt;&amp;#34;&lt;/span&gt;&lt;span class="k"&gt;$(&lt;/span&gt;security find-generic-password -s &lt;span class="s1"&gt;&amp;#39;Claude Code-credentials&amp;#39;&lt;/span&gt; -w&lt;span class="k"&gt;)&lt;/span&gt;&lt;span class="s2"&gt;&amp;#34;&lt;/span&gt;
&lt;/span&gt;&lt;/span&gt;&lt;span class="line"&gt;&lt;span class="cl"&gt;&lt;span class="o"&gt;}&lt;/span&gt;
&lt;/span&gt;&lt;/span&gt;&lt;span class="line"&gt;&lt;span class="cl"&gt;
&lt;/span&gt;&lt;/span&gt;&lt;span class="line"&gt;&lt;span class="cl"&gt;cc-use&lt;span class="o"&gt;()&lt;/span&gt; &lt;span class="o"&gt;{&lt;/span&gt; &lt;span class="c1"&gt;# cc-use personal&lt;/span&gt;
&lt;/span&gt;&lt;/span&gt;&lt;span class="line"&gt;&lt;span class="cl"&gt; security add-generic-password -U -a &lt;span class="s2"&gt;&amp;#34;&lt;/span&gt;&lt;span class="nv"&gt;$USER&lt;/span&gt;&lt;span class="s2"&gt;&amp;#34;&lt;/span&gt; -s &lt;span class="s2"&gt;&amp;#34;Claude Code-credentials&amp;#34;&lt;/span&gt; &lt;span class="se"&gt;\
&lt;/span&gt;&lt;/span&gt;&lt;/span&gt;&lt;span class="line"&gt;&lt;span class="cl"&gt; -w &lt;span class="s2"&gt;&amp;#34;&lt;/span&gt;&lt;span class="k"&gt;$(&lt;/span&gt;security find-generic-password -s &lt;span class="s2"&gt;&amp;#34;Claude Code-credentials-&lt;/span&gt;&lt;span class="nv"&gt;$1&lt;/span&gt;&lt;span class="s2"&gt;&amp;#34;&lt;/span&gt; -w&lt;span class="k"&gt;)&lt;/span&gt;&lt;span class="s2"&gt;&amp;#34;&lt;/span&gt;
&lt;/span&gt;&lt;/span&gt;&lt;span class="line"&gt;&lt;span class="cl"&gt;&lt;span class="o"&gt;}&lt;/span&gt;
&lt;/span&gt;&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;&lt;/div&gt;&lt;p&gt;Copy the live token into a named Keychain entry, and copy a named one back. That&amp;rsquo;s the whole trick, but it has two problems.&lt;/p&gt;
&lt;p&gt;The first is that OAuth tokens refresh. Claude Code quietly swaps in a fresh access token and writes it back to the Keychain, and the refresh token can rotate along with it, so a copy you saved yesterday can be dead today. The saved copy has to be kept fresh, which means saving before every switch and ideally all the time.&lt;/p&gt;
&lt;p&gt;The second is that I was only swapping the secret. &lt;code&gt;~/.claude.json&lt;/code&gt; still said I was the old account, so &lt;code&gt;/status&lt;/code&gt; showed the wrong email until I ran &lt;code&gt;/login&lt;/code&gt; anyway, which defeated the whole point.&lt;/p&gt;
&lt;h2 id="how-oof-switches-accounts"&gt;How oof switches accounts&lt;a class="heading-anchor" href="#how-oof-switches-accounts" aria-label="Link to this section"&gt;&lt;/a&gt;
&lt;/h2&gt;
&lt;p&gt;&lt;code&gt;oof&lt;/code&gt; fixes both problems and adds the comfort features I wanted once the basics worked. Each saved account gets:&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;a Keychain entry named &lt;code&gt;oof:&amp;lt;email&amp;gt;#&amp;lt;org&amp;gt;&lt;/code&gt; that holds its tokens. The org part matters because one email can belong to several organizations, a personal plan and a Team seat for example, and those are different logins.&lt;/li&gt;
&lt;li&gt;a JSON file in &lt;code&gt;~/.claude/accounts/&lt;/code&gt; with its &lt;code&gt;oauthAccount&lt;/code&gt; block. No secrets go on disk, only the label.&lt;/li&gt;
&lt;li&gt;an optional metadata file with a short name like &lt;code&gt;work&lt;/code&gt; and the plan, so I can type &lt;code&gt;oof work&lt;/code&gt; instead of an email.&lt;/li&gt;
&lt;/ul&gt;
&lt;p&gt;The switch itself is small. It reads the target&amp;rsquo;s tokens with &lt;code&gt;stored_secret&lt;/code&gt;, which looks in the account&amp;rsquo;s own Keychain entry, writes them into the slot Claude Code reads from, and then patches &lt;code&gt;oauthAccount&lt;/code&gt; in &lt;code&gt;~/.claude.json&lt;/code&gt;:&lt;/p&gt;
&lt;div class="highlight"&gt;&lt;pre tabindex="0" class="chroma"&gt;&lt;code class="language-zsh" data-lang="zsh"&gt;&lt;span class="line"&gt;&lt;span class="cl"&gt;switch_to&lt;span class="o"&gt;()&lt;/span&gt; &lt;span class="o"&gt;{&lt;/span&gt;
&lt;/span&gt;&lt;/span&gt;&lt;span class="line"&gt;&lt;span class="cl"&gt; &lt;span class="nb"&gt;local&lt;/span&gt; &lt;span class="nv"&gt;key&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="nv"&gt;$1&lt;/span&gt; secret tmp
&lt;/span&gt;&lt;/span&gt;&lt;span class="line"&gt;&lt;span class="cl"&gt; &lt;span class="nv"&gt;secret&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="k"&gt;$(&lt;/span&gt;stored_secret &lt;span class="s2"&gt;&amp;#34;&lt;/span&gt;&lt;span class="nv"&gt;$key&lt;/span&gt;&lt;span class="s2"&gt;&amp;#34;&lt;/span&gt;&lt;span class="k"&gt;)&lt;/span&gt; &lt;span class="se"&gt;\
&lt;/span&gt;&lt;/span&gt;&lt;/span&gt;&lt;span class="line"&gt;&lt;span class="cl"&gt; &lt;span class="o"&gt;||&lt;/span&gt; die &lt;span class="s2"&gt;&amp;#34;no saved credentials for &lt;/span&gt;&lt;span class="k"&gt;$(&lt;/span&gt;name_of &lt;span class="s2"&gt;&amp;#34;&lt;/span&gt;&lt;span class="nv"&gt;$key&lt;/span&gt;&lt;span class="s2"&gt;&amp;#34;&lt;/span&gt;&lt;span class="k"&gt;)&lt;/span&gt;&lt;span class="s2"&gt;. Run oof login to sign in again.&amp;#34;&lt;/span&gt;
&lt;/span&gt;&lt;/span&gt;&lt;span class="line"&gt;&lt;span class="cl"&gt; security add-generic-password -U -a &lt;span class="s2"&gt;&amp;#34;&lt;/span&gt;&lt;span class="nv"&gt;$USER&lt;/span&gt;&lt;span class="s2"&gt;&amp;#34;&lt;/span&gt; -s &lt;span class="s2"&gt;&amp;#34;&lt;/span&gt;&lt;span class="nv"&gt;$SLOT&lt;/span&gt;&lt;span class="s2"&gt;&amp;#34;&lt;/span&gt; -w &lt;span class="s2"&gt;&amp;#34;&lt;/span&gt;&lt;span class="nv"&gt;$secret&lt;/span&gt;&lt;span class="s2"&gt;&amp;#34;&lt;/span&gt; &lt;span class="se"&gt;\
&lt;/span&gt;&lt;/span&gt;&lt;/span&gt;&lt;span class="line"&gt;&lt;span class="cl"&gt; &lt;span class="o"&gt;||&lt;/span&gt; die &lt;span class="s2"&gt;&amp;#34;couldn&amp;#39;t write to the Keychain&amp;#34;&lt;/span&gt;
&lt;/span&gt;&lt;/span&gt;&lt;span class="line"&gt;&lt;span class="cl"&gt; &lt;span class="nv"&gt;tmp&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="k"&gt;$(&lt;/span&gt;mktemp &lt;span class="s2"&gt;&amp;#34;&lt;/span&gt;&lt;span class="nv"&gt;$CONFIG&lt;/span&gt;&lt;span class="s2"&gt;.oof.XXXXXX&amp;#34;&lt;/span&gt;&lt;span class="k"&gt;)&lt;/span&gt;
&lt;/span&gt;&lt;/span&gt;&lt;span class="line"&gt;&lt;span class="cl"&gt; jq --slurpfile acct &lt;span class="s2"&gt;&amp;#34;&lt;/span&gt;&lt;span class="nv"&gt;$STORE&lt;/span&gt;&lt;span class="s2"&gt;/&lt;/span&gt;&lt;span class="nv"&gt;$key&lt;/span&gt;&lt;span class="s2"&gt;.json&amp;#34;&lt;/span&gt; &lt;span class="s1"&gt;&amp;#39;.oauthAccount = $acct[0]&amp;#39;&lt;/span&gt; &lt;span class="s2"&gt;&amp;#34;&lt;/span&gt;&lt;span class="nv"&gt;$CONFIG&lt;/span&gt;&lt;span class="s2"&gt;&amp;#34;&lt;/span&gt; &amp;gt; &lt;span class="s2"&gt;&amp;#34;&lt;/span&gt;&lt;span class="nv"&gt;$tmp&lt;/span&gt;&lt;span class="s2"&gt;&amp;#34;&lt;/span&gt; &lt;span class="se"&gt;\
&lt;/span&gt;&lt;/span&gt;&lt;/span&gt;&lt;span class="line"&gt;&lt;span class="cl"&gt; &lt;span class="o"&gt;||&lt;/span&gt; &lt;span class="o"&gt;{&lt;/span&gt; rm -f &lt;span class="s2"&gt;&amp;#34;&lt;/span&gt;&lt;span class="nv"&gt;$tmp&lt;/span&gt;&lt;span class="s2"&gt;&amp;#34;&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt; die &lt;span class="s2"&gt;&amp;#34;couldn&amp;#39;t update &lt;/span&gt;&lt;span class="nv"&gt;$CONFIG&lt;/span&gt;&lt;span class="s2"&gt;&amp;#34;&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt; &lt;span class="o"&gt;}&lt;/span&gt;
&lt;/span&gt;&lt;/span&gt;&lt;span class="line"&gt;&lt;span class="cl"&gt; chmod &lt;span class="s2"&gt;&amp;#34;&lt;/span&gt;&lt;span class="k"&gt;$(&lt;/span&gt;stat -f %Lp &lt;span class="s2"&gt;&amp;#34;&lt;/span&gt;&lt;span class="nv"&gt;$CONFIG&lt;/span&gt;&lt;span class="s2"&gt;&amp;#34;&lt;/span&gt;&lt;span class="k"&gt;)&lt;/span&gt;&lt;span class="s2"&gt;&amp;#34;&lt;/span&gt; &lt;span class="s2"&gt;&amp;#34;&lt;/span&gt;&lt;span class="nv"&gt;$tmp&lt;/span&gt;&lt;span class="s2"&gt;&amp;#34;&lt;/span&gt;
&lt;/span&gt;&lt;/span&gt;&lt;span class="line"&gt;&lt;span class="cl"&gt; mv &lt;span class="s2"&gt;&amp;#34;&lt;/span&gt;&lt;span class="nv"&gt;$tmp&lt;/span&gt;&lt;span class="s2"&gt;&amp;#34;&lt;/span&gt; &lt;span class="s2"&gt;&amp;#34;&lt;/span&gt;&lt;span class="nv"&gt;$CONFIG&lt;/span&gt;&lt;span class="s2"&gt;&amp;#34;&lt;/span&gt;
&lt;/span&gt;&lt;/span&gt;&lt;span class="line"&gt;&lt;span class="cl"&gt;&lt;span class="o"&gt;}&lt;/span&gt;
&lt;/span&gt;&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;&lt;/div&gt;&lt;p&gt;&lt;code&gt;~/.claude.json&lt;/code&gt; holds a lot more than the login, including per-project settings and MCP servers, so I never edit it in place. &lt;code&gt;jq&lt;/code&gt; writes a temp file next to it, the temp file gets the original&amp;rsquo;s permissions, and &lt;code&gt;mv&lt;/code&gt; swaps it in with a single rename. If anything fails halfway, the real file is untouched.&lt;/p&gt;
&lt;p&gt;Before any of that runs, &lt;code&gt;oof&lt;/code&gt; saves the current login. That&amp;rsquo;s the fix for stale tokens: every switch starts by copying the live, freshly refreshed tokens into the current account&amp;rsquo;s entry, so the account you&amp;rsquo;re leaving is always stored in its newest state. Then it &lt;code&gt;exec&lt;/code&gt;s &lt;code&gt;claude -c&lt;/code&gt; and the conversation picks up where it left off.&lt;/p&gt;
&lt;p&gt;Picking the target is meant to need as little typing as possible:&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;&lt;code&gt;oof&lt;/code&gt; on its own rotates to the next saved account.&lt;/li&gt;
&lt;li&gt;&lt;code&gt;oof work&lt;/code&gt; first looks for an exact short name or email, and if nothing matches it tries a substring, so &lt;code&gt;oof pers&lt;/code&gt; is enough. If a substring matches two accounts it refuses and names both, because guessing wrong here means quietly working as the wrong person.&lt;/li&gt;
&lt;li&gt;&lt;code&gt;oof pick&lt;/code&gt; opens a menu.&lt;/li&gt;
&lt;li&gt;Anything after &lt;code&gt;--&lt;/code&gt; goes to &lt;code&gt;claude&lt;/code&gt; in place of &lt;code&gt;-c&lt;/code&gt;, so &lt;code&gt;oof pick -- -r&lt;/code&gt; lets me choose an account and then choose a conversation.&lt;/li&gt;
&lt;/ul&gt;
&lt;h2 id="saving-logins-automatically-with-claude-code-hooks"&gt;Saving logins automatically with Claude Code hooks&lt;a class="heading-anchor" href="#saving-logins-automatically-with-claude-code-hooks" aria-label="Link to this section"&gt;&lt;/a&gt;
&lt;/h2&gt;
&lt;p&gt;Saving before every switch covers switches made with &lt;code&gt;oof&lt;/code&gt;. It doesn&amp;rsquo;t cover running &lt;code&gt;/login&lt;/code&gt; inside Claude Code by hand, which is how a new account usually gets added. For that I leaned on Claude Code&amp;rsquo;s hooks. Two hooks in &lt;code&gt;~/.claude/settings.json&lt;/code&gt; run &lt;code&gt;oof save -q&lt;/code&gt; in the background when a session starts and after every reply:&lt;/p&gt;
&lt;div class="highlight"&gt;&lt;pre tabindex="0" class="chroma"&gt;&lt;code class="language-json" data-lang="json"&gt;&lt;span class="line"&gt;&lt;span class="cl"&gt;&lt;span class="p"&gt;{&lt;/span&gt;
&lt;/span&gt;&lt;/span&gt;&lt;span class="line"&gt;&lt;span class="cl"&gt; &lt;span class="nt"&gt;&amp;#34;hooks&amp;#34;&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt;
&lt;/span&gt;&lt;/span&gt;&lt;span class="line"&gt;&lt;span class="cl"&gt; &lt;span class="nt"&gt;&amp;#34;SessionStart&amp;#34;&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="p"&gt;[{&lt;/span&gt; &lt;span class="nt"&gt;&amp;#34;hooks&amp;#34;&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="p"&gt;[{&lt;/span&gt; &lt;span class="nt"&gt;&amp;#34;type&amp;#34;&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="s2"&gt;&amp;#34;command&amp;#34;&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="nt"&gt;&amp;#34;command&amp;#34;&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="s2"&gt;&amp;#34;\&amp;#34;$HOME/.local/bin/oof\&amp;#34; save -q || true&amp;#34;&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="nt"&gt;&amp;#34;timeout&amp;#34;&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="mi"&gt;15&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="nt"&gt;&amp;#34;async&amp;#34;&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="kc"&gt;true&lt;/span&gt; &lt;span class="p"&gt;}]&lt;/span&gt; &lt;span class="p"&gt;}],&lt;/span&gt;
&lt;/span&gt;&lt;/span&gt;&lt;span class="line"&gt;&lt;span class="cl"&gt; &lt;span class="nt"&gt;&amp;#34;Stop&amp;#34;&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="p"&gt;[{&lt;/span&gt; &lt;span class="nt"&gt;&amp;#34;hooks&amp;#34;&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="p"&gt;[{&lt;/span&gt; &lt;span class="nt"&gt;&amp;#34;type&amp;#34;&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="s2"&gt;&amp;#34;command&amp;#34;&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="nt"&gt;&amp;#34;command&amp;#34;&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="s2"&gt;&amp;#34;\&amp;#34;$HOME/.local/bin/oof\&amp;#34; save -q || true&amp;#34;&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="nt"&gt;&amp;#34;timeout&amp;#34;&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="mi"&gt;15&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="nt"&gt;&amp;#34;async&amp;#34;&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="kc"&gt;true&lt;/span&gt; &lt;span class="p"&gt;}]&lt;/span&gt; &lt;span class="p"&gt;}]&lt;/span&gt;
&lt;/span&gt;&lt;/span&gt;&lt;span class="line"&gt;&lt;span class="cl"&gt; &lt;span class="p"&gt;}&lt;/span&gt;
&lt;/span&gt;&lt;/span&gt;&lt;span class="line"&gt;&lt;span class="cl"&gt;&lt;span class="p"&gt;}&lt;/span&gt;
&lt;/span&gt;&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;&lt;/div&gt;&lt;p&gt;Now any login gets picked up after the next message, and the tokens on file are never more than one reply old. Since that save runs constantly, it compares the live secret with the stored one first and only writes to the Keychain when they differ, otherwise it would rewrite a Keychain item every few seconds for nothing. The &lt;code&gt;|| true&lt;/code&gt; and &lt;code&gt;async&lt;/code&gt; make sure a hook failure never gets in the way of a reply.&lt;/p&gt;
&lt;p&gt;&lt;code&gt;oof login&lt;/code&gt; does the same from a normal terminal. It saves the current account, runs &lt;code&gt;claude auth login&lt;/code&gt;, saves the new one and suggests a short name for it.&lt;/p&gt;
&lt;h2 id="making-it-pleasant"&gt;Making it pleasant&lt;a class="heading-anchor" href="#making-it-pleasant" aria-label="Link to this section"&gt;&lt;/a&gt;
&lt;/h2&gt;
&lt;p&gt;Once it worked I gave it some polish, because a tool you reach for in the middle of real work should feel good to run.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;The banner.&lt;/strong&gt; &lt;code&gt;oof help&lt;/code&gt; opens with a big OOF in block letters that fades through Claude&amp;rsquo;s orange from top to bottom. The letters are solid blocks with box-drawing characters as a drop shadow. I wanted the blocks colored and the shadow dim, and zsh can do that with one substitution per line:&lt;/p&gt;
&lt;div class="highlight"&gt;&lt;pre tabindex="0" class="chroma"&gt;&lt;code class="language-zsh" data-lang="zsh"&gt;&lt;span class="line"&gt;&lt;span class="cl"&gt;&lt;span class="nv"&gt;line&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="si"&gt;${&lt;/span&gt;&lt;span class="nv"&gt;art&lt;/span&gt;&lt;span class="p"&gt;[i]//(#m)[╔╗╚╝═║]##/&lt;/span&gt;&lt;span class="si"&gt;${&lt;/span&gt;&lt;span class="nv"&gt;R&lt;/span&gt;&lt;span class="si"&gt;}${&lt;/span&gt;&lt;span class="nv"&gt;D&lt;/span&gt;&lt;span class="si"&gt;}${&lt;/span&gt;&lt;span class="nv"&gt;MATCH&lt;/span&gt;&lt;span class="si"&gt;}${&lt;/span&gt;&lt;span class="nv"&gt;R&lt;/span&gt;&lt;span class="si"&gt;}${&lt;/span&gt;&lt;span class="nv"&gt;grad&lt;/span&gt;&lt;span class="p"&gt;[i]&lt;/span&gt;&lt;span class="si"&gt;}}&lt;/span&gt;
&lt;/span&gt;&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;&lt;/div&gt;&lt;p&gt;With &lt;code&gt;extended_glob&lt;/code&gt; on, &lt;code&gt;(#m)&lt;/code&gt; captures each run of shadow characters into &lt;code&gt;$MATCH&lt;/code&gt;, and the replacement wraps it in &amp;ldquo;reset, dim, the shadow, then back to this row&amp;rsquo;s gradient color&amp;rdquo;. It uses truecolor when the terminal supports it, falls back to the 256-color palette otherwise, and turns color off entirely for pipes, &lt;code&gt;NO_COLOR&lt;/code&gt; and dumb terminals.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;The menu.&lt;/strong&gt; &lt;code&gt;oof pick&lt;/code&gt; is an arrow-key menu in plain zsh, with no fzf or other dependency. It opens &lt;code&gt;/dev/tty&lt;/code&gt; on its own file descriptor, so the menu keeps working while stdout is captured, which matters because the chosen account is printed back to the caller. Then it puts the terminal into raw mode with &lt;code&gt;stty -icanon -echo&lt;/code&gt; and reads one key at a time. The fiddly part is the Escape key. An arrow key arrives as Escape followed by &lt;code&gt;[A&lt;/code&gt;, while plain Escape arrives alone, so after an Escape the menu waits 50 milliseconds for more bytes. If &lt;code&gt;[A&lt;/code&gt; or &lt;code&gt;[B&lt;/code&gt; shows up it moves the selection, and if nothing comes it treats it as Escape and cancels. The menu redraws in place by moving the cursor up and clearing each line, a trap restores the terminal if you hit Ctrl+C partway through, and it starts on the next account rather than the current one, since you opened it to leave.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;The list.&lt;/strong&gt; &lt;code&gt;oof ls&lt;/code&gt; shows each account&amp;rsquo;s short name, email, plan and when it was last active, with a green dot on the current one. The plan comes from the stored token data, where the subscription type and rate tier turn into labels like &amp;ldquo;Max 20x&amp;rdquo; or &amp;ldquo;Team&amp;rdquo;. &amp;ldquo;Last active&amp;rdquo; is simply the modification time of the account&amp;rsquo;s file, which the hooks refresh on every save.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Completion.&lt;/strong&gt; Pressing Tab after &lt;code&gt;oof&lt;/code&gt; lists the saved accounts, with each one&amp;rsquo;s email, plan and whether it&amp;rsquo;s active, next to the commands. A hidden &lt;code&gt;oof __accounts&lt;/code&gt; command prints one &lt;code&gt;value:description&lt;/code&gt; line per account, and the completion function hands those to zsh&amp;rsquo;s &lt;code&gt;_describe&lt;/code&gt;. The completion script itself comes from &lt;code&gt;oof completion zsh&lt;/code&gt;, so it can&amp;rsquo;t drift out of sync with the commands. One gotcha worth knowing: if your &lt;code&gt;.zshrc&lt;/code&gt; runs &lt;code&gt;compinit -C&lt;/code&gt;, zsh trusts its cached dump and won&amp;rsquo;t notice a new completion file until you delete &lt;code&gt;~/.zcompdump&lt;/code&gt;.&lt;/p&gt;
&lt;h2 id="is-it-safe-with-several-sessions-open"&gt;Is it safe with several sessions open?&lt;a class="heading-anchor" href="#is-it-safe-with-several-sessions-open" aria-label="Link to this section"&gt;&lt;/a&gt;
&lt;/h2&gt;
&lt;p&gt;I usually have a handful of Claude Code sessions open in cmux, so this was my main worry. Say I switch to account B while an older session still has account A in memory. If that session refreshed its token, it might write A&amp;rsquo;s tokens back over B&amp;rsquo;s, and then my hook would save A&amp;rsquo;s tokens under B&amp;rsquo;s name.&lt;/p&gt;
&lt;p&gt;Before adding locks and timers I looked at what Claude Code actually does when it refreshes, and it already handles this. Before refreshing, it re-reads what&amp;rsquo;s stored, and if the stored token isn&amp;rsquo;t the one it was holding, it adopts the stored one instead of refreshing its own. When it does refresh, it only writes the result back if the stored token is still the one it started from. An old session can&amp;rsquo;t overwrite a newer login, so the scary version of the race doesn&amp;rsquo;t happen.&lt;/p&gt;
&lt;p&gt;There are two smaller rough edges I know about and haven&amp;rsquo;t fixed yet:&lt;/p&gt;
&lt;ol&gt;
&lt;li&gt;&lt;strong&gt;The token and the label are two separate writes.&lt;/strong&gt; If a hook&amp;rsquo;s save lands in the few milliseconds between &lt;code&gt;oof&lt;/code&gt; writing the Keychain and writing &lt;code&gt;~/.claude.json&lt;/code&gt;, it pairs the new account&amp;rsquo;s token with the old account&amp;rsquo;s name. A &lt;code&gt;/login&lt;/code&gt; in another window has a wider version of the same gap, since it fetches the profile over the network between its two writes. The proper fix is to hold Claude Code&amp;rsquo;s own lock files during a switch and have &lt;code&gt;save&lt;/code&gt; stand down while a switch is in progress. Until then, &lt;code&gt;oof&lt;/code&gt; warns when other &lt;code&gt;claude&lt;/code&gt; processes are running so I can restart them, and if a pairing ever goes wrong, one &lt;code&gt;/login&lt;/code&gt; repairs it on the next reply.&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;The secret briefly appears in a process argument.&lt;/strong&gt; &lt;code&gt;security add-generic-password -w &amp;quot;$secret&amp;quot;&lt;/code&gt; puts the token on the command line, where &lt;code&gt;ps&lt;/code&gt; can see it for a moment. Claude Code avoids this by feeding &lt;code&gt;security -i&lt;/code&gt; its commands on stdin. On a single-user laptop the risk is small, but it&amp;rsquo;s the next thing I&amp;rsquo;d change.&lt;/li&gt;
&lt;/ol&gt;
&lt;h2 id="install-oof"&gt;Install oof&lt;a class="heading-anchor" href="#install-oof" aria-label="Link to this section"&gt;&lt;/a&gt;
&lt;/h2&gt;
&lt;p&gt;You need macOS, zsh, Claude Code and &lt;code&gt;jq&lt;/code&gt; (&lt;code&gt;brew install jq&lt;/code&gt;). It&amp;rsquo;s one file, so read it before you run it:&lt;/p&gt;
&lt;div class="highlight"&gt;&lt;pre tabindex="0" class="chroma"&gt;&lt;code class="language-zsh" data-lang="zsh"&gt;&lt;span class="line"&gt;&lt;span class="cl"&gt;mkdir -p ~/.local/bin
&lt;/span&gt;&lt;/span&gt;&lt;span class="line"&gt;&lt;span class="cl"&gt;curl -fsSL https://gist.githubusercontent.com/ankushKun/cd27fb07b0f06be99d4181b40309103d/raw/oof.sh -o ~/.local/bin/oof
&lt;/span&gt;&lt;/span&gt;&lt;span class="line"&gt;&lt;span class="cl"&gt;chmod +x ~/.local/bin/oof
&lt;/span&gt;&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;&lt;/div&gt;&lt;p&gt;Make sure &lt;code&gt;~/.local/bin&lt;/code&gt; is on your &lt;code&gt;PATH&lt;/code&gt;, then:&lt;/p&gt;
&lt;ol&gt;
&lt;li&gt;Add the two hooks above to &lt;code&gt;~/.claude/settings.json&lt;/code&gt;, merging them with any hooks you already have. You can confirm they&amp;rsquo;re active with &lt;code&gt;/hooks&lt;/code&gt; inside Claude Code.&lt;/li&gt;
&lt;li&gt;Install tab completion with &lt;code&gt;oof completion zsh &amp;gt; &amp;quot;$(brew --prefix)/share/zsh/site-functions/_oof&amp;quot;&lt;/code&gt;, then run &lt;code&gt;rm -f ~/.zcompdump&lt;/code&gt; and open a new terminal.&lt;/li&gt;
&lt;li&gt;Run &lt;code&gt;oof login&lt;/code&gt; once for each account you want to add, and give each one a short name with &lt;code&gt;oof name you@company.com work&lt;/code&gt;.&lt;/li&gt;
&lt;/ol&gt;
&lt;p&gt;After that, &lt;code&gt;oof&lt;/code&gt; rotates to the next account, &lt;code&gt;oof work&lt;/code&gt; jumps to a named one, &lt;code&gt;oof pick&lt;/code&gt; shows the menu, and &lt;code&gt;oof ls&lt;/code&gt; lists everything. The first time it touches the Keychain, macOS may ask for permission, and &amp;ldquo;Always Allow&amp;rdquo; stops it asking again.&lt;/p&gt;
&lt;h2 id="a-note-on-limits"&gt;A note on limits&lt;a class="heading-anchor" href="#a-note-on-limits" aria-label="Link to this section"&gt;&lt;/a&gt;
&lt;/h2&gt;
&lt;p&gt;The name is a joke about hitting a limit, but &lt;code&gt;oof&lt;/code&gt; is for switching between accounts you actually have, like a personal plan and a seat on your company&amp;rsquo;s Team plan. Each account keeps its own limits, and switching doesn&amp;rsquo;t change them. If what you keep running into is the usage limit on a single plan, the supported routes are extra usage, a bigger plan or an API key. Before you set up a stack of personal subscriptions just to rotate through them, read Anthropic&amp;rsquo;s terms first.&lt;/p&gt;
&lt;h2 id="wrapping-up"&gt;Wrapping up&lt;a class="heading-anchor" href="#wrapping-up" aria-label="Link to this section"&gt;&lt;/a&gt;
&lt;/h2&gt;
&lt;p&gt;The whole thing is under 600 lines of zsh, and most of that is help text, colors and the menu. The core idea fits in a paragraph. Conversations live on disk and not in your account, and the login is a Keychain item plus a label. Keep a fresh copy of each account&amp;rsquo;s item, swap both pieces together, and switching becomes one command that drops you back into the same conversation.&lt;/p&gt;
&lt;p&gt;I built it in one sitting with Claude Code, against Claude Code 2.1.296. For testing, the &lt;code&gt;OOF_CONFIG&lt;/code&gt; and &lt;code&gt;OOF_STORE&lt;/code&gt; variables point the script at a throwaway config and account store, so the menus, naming and completion could be exercised against fake accounts without touching real logins. That&amp;rsquo;s also how the screenshot at the top was made. The script is &lt;a href="https://gist.github.com/ankushKun/cd27fb07b0f06be99d4181b40309103d" data-external&gt;on GitHub Gist&lt;/a&gt; if you want to read it, use it or adapt it for Linux. It started life as &lt;code&gt;ccs&lt;/code&gt;, and logins saved by that version still work after the rename.&lt;/p&gt;</content:encoded></item><item><title>Clef beat Jev on Jev's home turf</title><link>https://ankush.one/blogs/clef-vs-jev-benchmark/</link><pubDate>Mon, 05 Oct 2026 00:00:00 +0000</pubDate><author>ankush4singh@gmail.com (Ankush Singh)</author><guid>https://ankush.one/blogs/clef-vs-jev-benchmark/</guid><category>ai-agents</category><category>machine-learning</category><category>cloudflare</category><category>construct</category><description>We ran Cloudflare&amp;rsquo;s Clef and Clef-flash against Jev on five production AI agent decisions: 1,053 calls, accuracy, determinism, drift, cost, and what we are switching.</description><content:encoded>&lt;p&gt;When Cloudflare released Clef on October 1, its launch post carried a table of ten public benchmarks, seven of them led by a Clef model. The Hacker News thread passed 600 points in a day, and one of the first replies was six words long: &amp;ldquo;Public benchmarks are easy to cheat&amp;rdquo; (&lt;a href="https://traictory.com/news/2026-10-03-cloudflare-clef-decision-models" data-external&gt;Traictory&lt;/a&gt;, &lt;a href="https://news.ycombinator.com/item?id=49923692" data-external&gt;Hacker News&lt;/a&gt;).&lt;/p&gt;
&lt;p&gt;Fair. So we ran the test nobody could have trained for. Construct&amp;rsquo;s agent has about 40 small decisions written for a decision model, and every one of them was written around Jev: the prompts, the questions, and the thresholds, all built against Jev over the last two weeks. That is Jev&amp;rsquo;s home turf. We pointed the same requests at Clef and Clef-flash without changing a line, and scored all three against labelled cases.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Clef is Cloudflare&amp;rsquo;s open-weight decision model, a drop-in alternative to TypeSafe&amp;rsquo;s Jev. On 117 labelled cases across five of our production decisions, Clef reached the right outcome on 92.3% of calls using thresholds set for Jev, against 88.6% for Jev and 82.1% for Clef-flash. Both Clef models returned the identical probability on every repeated call, and Jev did not.&lt;/strong&gt;&lt;/p&gt;
&lt;p&gt;We build Construct, an AI employee with its own computer. Every number below comes from our own runs on &lt;strong&gt;October 4, 2026&lt;/strong&gt; (UTC): 117 cases, three calls each, three models, 1,053 calls in the main run, plus about 800 more in follow-up controls. Facts about the models themselves come from Cloudflare&amp;rsquo;s and TypeSafe&amp;rsquo;s documentation, checked the same day. The full method, confidence intervals and per-case results are in the paper: &lt;a href="https://construct.computer/blog/clef-vs-jev-benchmark/compatible-apis-incompatible-thresholds.pdf" data-external&gt;Compatible APIs, Incompatible Thresholds: A Drop-In Replacement Study of Clef and Jev on Decision Specifications from a Production Agent&lt;/a&gt; (PDF, 13 pages).&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;&lt;strong&gt;Accuracy, drop-in:&lt;/strong&gt; Clef 92.3%, Jev 88.6%, Clef-flash 82.1%, all on thresholds set for Jev. The gap between Clef and Jev is not statistically significant.&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Accuracy, own thresholds:&lt;/strong&gt; Clef 93.2%, Jev 90.9%, Clef-flash 90.6%, with each threshold chosen on held-out cases.&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Repeatability:&lt;/strong&gt; Clef and Clef-flash gave the same probability on all 351 calls each. The probability Jev&amp;rsquo;s decisions act on moved by up to 0.10, and 2 cases changed verdict.&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Code changed to swap models:&lt;/strong&gt; one string.&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Failed calls:&lt;/strong&gt; 0 of 1,053.&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Cost:&lt;/strong&gt; $23 per million decisions on Jev, $35 on Clef-flash, $93 on Clef. The main run cost about five cents.&lt;/li&gt;
&lt;/ul&gt;
&lt;h2 id="what-are-clef-and-clef-flash"&gt;What are Clef and Clef-flash?&lt;a class="heading-anchor" href="#what-are-clef-and-clef-flash" aria-label="Link to this section"&gt;&lt;/a&gt;
&lt;/h2&gt;
&lt;p&gt;Clef and Clef-flash are decision models from Cloudflare: you send a state and a set of typed questions, and they return a probability for each allowed answer instead of generating text. Clef is a 27 billion parameter model built on Qwen3.8, and Clef-flash is a 9 billion parameter model built on Qwen3.5. Both are released under Apache 2.0 with weights on Hugging Face, both run on Workers AI, and both accept the same request format as Jev (&lt;a href="https://blog.cloudflare.com/clef-decision-models/" data-external&gt;Cloudflare&lt;/a&gt;).&lt;/p&gt;
&lt;p&gt;If decision models are new to you: they are the fast, cheap judgment calls around a language model. Is this the same person? Does this email need a reply? Should this tool call wait for a human? TypeSafe calls the category System One models, after Kahneman&amp;rsquo;s fast, intuitive System 1. Our &lt;a href="https://construct.computer/blog/jev-ai-agents/" data-external&gt;first post on Jev&lt;/a&gt; explains the idea and how we put one inside an AI agent&amp;rsquo;s memory.&lt;/p&gt;
&lt;table&gt;
	&lt;thead&gt;
			&lt;tr&gt;
					&lt;th&gt;&lt;/th&gt;
					&lt;th&gt;Jev 1.13&lt;/th&gt;
					&lt;th&gt;Clef&lt;/th&gt;
					&lt;th&gt;Clef-flash&lt;/th&gt;
			&lt;/tr&gt;
	&lt;/thead&gt;
	&lt;tbody&gt;
			&lt;tr&gt;
					&lt;td&gt;Maker&lt;/td&gt;
					&lt;td&gt;TypeSafe AI&lt;/td&gt;
					&lt;td&gt;Cloudflare&lt;/td&gt;
					&lt;td&gt;Cloudflare&lt;/td&gt;
			&lt;/tr&gt;
			&lt;tr&gt;
					&lt;td&gt;Size&lt;/td&gt;
					&lt;td&gt;Not published&lt;/td&gt;
					&lt;td&gt;27B parameters&lt;/td&gt;
					&lt;td&gt;9B parameters&lt;/td&gt;
			&lt;/tr&gt;
			&lt;tr&gt;
					&lt;td&gt;Weights&lt;/td&gt;
					&lt;td&gt;Closed&lt;/td&gt;
					&lt;td&gt;Open, Apache 2.0&lt;/td&gt;
					&lt;td&gt;Open, Apache 2.0&lt;/td&gt;
			&lt;/tr&gt;
			&lt;tr&gt;
					&lt;td&gt;Question types&lt;/td&gt;
					&lt;td&gt;Choice, Score, Noul&lt;/td&gt;
					&lt;td&gt;Same&lt;/td&gt;
					&lt;td&gt;Same&lt;/td&gt;
			&lt;/tr&gt;
			&lt;tr&gt;
					&lt;td&gt;Input&lt;/td&gt;
					&lt;td&gt;Text and JSON&lt;/td&gt;
					&lt;td&gt;Text, JSON and images&lt;/td&gt;
					&lt;td&gt;Text, JSON and images&lt;/td&gt;
			&lt;/tr&gt;
			&lt;tr&gt;
					&lt;td&gt;Context&lt;/td&gt;
					&lt;td&gt;64k tokens per request&lt;/td&gt;
					&lt;td&gt;64k tokens&lt;/td&gt;
					&lt;td&gt;64k tokens&lt;/td&gt;
			&lt;/tr&gt;
			&lt;tr&gt;
					&lt;td&gt;Price per million input tokens&lt;/td&gt;
					&lt;td&gt;$0.042&lt;/td&gt;
					&lt;td&gt;$0.24&lt;/td&gt;
					&lt;td&gt;$0.09&lt;/td&gt;
			&lt;/tr&gt;
			&lt;tr&gt;
					&lt;td&gt;Output tokens&lt;/td&gt;
					&lt;td&gt;Free&lt;/td&gt;
					&lt;td&gt;Free&lt;/td&gt;
					&lt;td&gt;Free&lt;/td&gt;
			&lt;/tr&gt;
			&lt;tr&gt;
					&lt;td&gt;Workers AI model ID&lt;/td&gt;
					&lt;td&gt;&lt;code&gt;typesafe/jev&lt;/code&gt;&lt;/td&gt;
					&lt;td&gt;&lt;code&gt;@cf/cloudflare/clef&lt;/code&gt;&lt;/td&gt;
					&lt;td&gt;&lt;code&gt;@cf/cloudflare/clef-flash&lt;/code&gt;&lt;/td&gt;
			&lt;/tr&gt;
	&lt;/tbody&gt;
&lt;/table&gt;
&lt;p&gt;&lt;em&gt;Jev, Clef and Clef-flash compared, October 2026&lt;/em&gt;&lt;/p&gt;
&lt;p&gt;Two details in that table matter later. Clef reads images, which Jev does not today. And the prices are per input token, so the model that counts fewer tokens for the same request closes some of the gap.&lt;/p&gt;
&lt;h2 id="how-we-benchmarked-clef-against-jev"&gt;How we benchmarked Clef against Jev&lt;a class="heading-anchor" href="#how-we-benchmarked-clef-against-jev" aria-label="Link to this section"&gt;&lt;/a&gt;
&lt;/h2&gt;
&lt;p&gt;We did not write a new test for Clef. We took five decisions that already exist in Construct&amp;rsquo;s code and ran each one&amp;rsquo;s real request builder and real interpreter, the same functions production calls.&lt;/p&gt;
&lt;table&gt;
	&lt;thead&gt;
			&lt;tr&gt;
					&lt;th&gt;Decision&lt;/th&gt;
					&lt;th&gt;What it decides&lt;/th&gt;
					&lt;th&gt;Cases&lt;/th&gt;
					&lt;th&gt;Threshold today&lt;/th&gt;
			&lt;/tr&gt;
	&lt;/thead&gt;
	&lt;tbody&gt;
			&lt;tr&gt;
					&lt;td&gt;Entity match&lt;/td&gt;
					&lt;td&gt;Is this newly mentioned person, project or company one the agent already knows?&lt;/td&gt;
					&lt;td&gt;30&lt;/td&gt;
					&lt;td&gt;0.70&lt;/td&gt;
			&lt;/tr&gt;
			&lt;tr&gt;
					&lt;td&gt;Email triage&lt;/td&gt;
					&lt;td&gt;Newsletter, receipt, spam, request, or keep?&lt;/td&gt;
					&lt;td&gt;15&lt;/td&gt;
					&lt;td&gt;0.80&lt;/td&gt;
			&lt;/tr&gt;
			&lt;tr&gt;
					&lt;td&gt;Memory extract gate&lt;/td&gt;
					&lt;td&gt;Does this conversation hold anything worth remembering?&lt;/td&gt;
					&lt;td&gt;24&lt;/td&gt;
					&lt;td&gt;0.05&lt;/td&gt;
			&lt;/tr&gt;
			&lt;tr&gt;
					&lt;td&gt;Notification priority&lt;/td&gt;
					&lt;td&gt;Does this automation notice deserve a toast, or can it wait?&lt;/td&gt;
					&lt;td&gt;24&lt;/td&gt;
					&lt;td&gt;0.85&lt;/td&gt;
			&lt;/tr&gt;
			&lt;tr&gt;
					&lt;td&gt;Tool-call risk&lt;/td&gt;
					&lt;td&gt;Should this tool call wait for the user&amp;rsquo;s go-ahead?&lt;/td&gt;
					&lt;td&gt;24&lt;/td&gt;
					&lt;td&gt;0.80&lt;/td&gt;
			&lt;/tr&gt;
	&lt;/tbody&gt;
&lt;/table&gt;
&lt;p&gt;&lt;em&gt;The five production decisions in the benchmark&lt;/em&gt;&lt;/p&gt;
&lt;p&gt;Entity match is live in production. The other four are built and switched off, waiting on shadow data. Their thresholds were set by judgment with Jev in mind and have never been fitted to data, for Jev or anything else.&lt;/p&gt;
&lt;p&gt;Swapping the model was the least interesting part of the work. This is the whole change:&lt;/p&gt;
&lt;div class="highlight"&gt;&lt;pre tabindex="0" class="chroma"&gt;&lt;code class="language-ts" data-lang="ts"&gt;&lt;span class="line"&gt;&lt;span class="cl"&gt;&lt;span class="c1"&gt;// Before
&lt;/span&gt;&lt;/span&gt;&lt;/span&gt;&lt;span class="line"&gt;&lt;span class="cl"&gt;&lt;span class="k"&gt;await&lt;/span&gt; &lt;span class="nx"&gt;env&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nx"&gt;AI&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nx"&gt;run&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="s2"&gt;&amp;#34;typesafe/jev&amp;#34;&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt; &lt;span class="nx"&gt;state&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="nx"&gt;questions&lt;/span&gt; &lt;span class="p"&gt;},&lt;/span&gt; &lt;span class="nx"&gt;options&lt;/span&gt;&lt;span class="p"&gt;);&lt;/span&gt;
&lt;/span&gt;&lt;/span&gt;&lt;span class="line"&gt;&lt;span class="cl"&gt;
&lt;/span&gt;&lt;/span&gt;&lt;span class="line"&gt;&lt;span class="cl"&gt;&lt;span class="c1"&gt;// After
&lt;/span&gt;&lt;/span&gt;&lt;/span&gt;&lt;span class="line"&gt;&lt;span class="cl"&gt;&lt;span class="k"&gt;await&lt;/span&gt; &lt;span class="nx"&gt;env&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nx"&gt;AI&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nx"&gt;run&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="s2"&gt;&amp;#34;@cf/cloudflare/clef&amp;#34;&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt; &lt;span class="nx"&gt;state&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="nx"&gt;questions&lt;/span&gt; &lt;span class="p"&gt;},&lt;/span&gt; &lt;span class="nx"&gt;options&lt;/span&gt;&lt;span class="p"&gt;);&lt;/span&gt;
&lt;/span&gt;&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;&lt;/div&gt;&lt;p&gt;Cloudflare says Clef is &amp;ldquo;fully Jev-API compatible&amp;rdquo;. On our requests that held: Choice, Score and Noul questions, JSON state, and the response shape all worked unchanged, and all 702 Clef calls returned valid answers.&lt;/p&gt;
&lt;p&gt;The rest of the method, briefly:&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;&lt;strong&gt;Cases:&lt;/strong&gt; 117, all synthetic, weighted toward the edges of each decision. We drafted and labelled them with an AI assistant (Claude) before any of the three models saw them. No label came from Jev or Clef, and the labels have not yet been independently reviewed by a second person.&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Runs:&lt;/strong&gt; each case three times per model, one call at a time, models interleaved so they shared the same network conditions, with the gateway cache skipped.&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Scoring:&lt;/strong&gt; a call is correct when the production interpreter reaches the labelled outcome at the production threshold. We also report each model&amp;rsquo;s best threshold, found on the same cases, which flatters all three equally.&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Route:&lt;/strong&gt; Cloudflare&amp;rsquo;s REST API through our AI Gateway, from a laptop. Fine for accuracy. Wrong for latency, which gets its own section.&lt;/li&gt;
&lt;/ul&gt;
&lt;h2 id="clef-vs-jev-accuracy-the-results"&gt;Clef vs Jev accuracy: the results&lt;a class="heading-anchor" href="#clef-vs-jev-accuracy-the-results" aria-label="Link to this section"&gt;&lt;/a&gt;
&lt;/h2&gt;
&lt;figure class="doc-figure"&gt;&lt;img src="https://ankush.one/media/construct/clef-vs-jev-benchmark/01-accuracy-overall.webp" alt="Bar chart of overall accuracy. With thresholds set for Jev: Jev 88.6%, Clef 92.3%, Clef-flash 82.1%. With each model’s own best threshold: Jev 93.7%, Clef 96.6%, Clef-flash 94.9%." title="Overall accuracy across five production decisions, 351 calls per model. Left: the thresholds configured today, set for Jev. Right: each model&amp;#39;s own best threshold, found in-sample." width="2000" height="1125" loading="lazy" decoding="async"&gt;&lt;figcaption&gt;Overall accuracy across five production decisions, 351 calls per model. Left: the thresholds configured today, set for Jev. Right: each model&amp;#39;s own best threshold, found in-sample.&lt;/figcaption&gt;&lt;/figure&gt;&lt;p&gt;With no retuning at all, Clef had the highest score of the three, and the highest or equal-highest on four of the five decisions, on thresholds that were set for a different model.&lt;/p&gt;
&lt;table&gt;
	&lt;thead&gt;
			&lt;tr&gt;
					&lt;th&gt;Decision&lt;/th&gt;
					&lt;th&gt;Jev 1.13&lt;/th&gt;
					&lt;th&gt;Clef&lt;/th&gt;
					&lt;th&gt;Clef-flash&lt;/th&gt;
			&lt;/tr&gt;
	&lt;/thead&gt;
	&lt;tbody&gt;
			&lt;tr&gt;
					&lt;td&gt;Entity match&lt;/td&gt;
					&lt;td&gt;72.2% (82.2%)&lt;/td&gt;
					&lt;td&gt;&lt;strong&gt;86.7%&lt;/strong&gt; (93.3%)&lt;/td&gt;
					&lt;td&gt;80.0% (86.7%)&lt;/td&gt;
			&lt;/tr&gt;
			&lt;tr&gt;
					&lt;td&gt;Email triage&lt;/td&gt;
					&lt;td&gt;&lt;strong&gt;93.3%&lt;/strong&gt; (93.3%)&lt;/td&gt;
					&lt;td&gt;86.7% (93.3%)&lt;/td&gt;
					&lt;td&gt;73.3% (86.7%)&lt;/td&gt;
			&lt;/tr&gt;
			&lt;tr&gt;
					&lt;td&gt;Memory extract gate&lt;/td&gt;
					&lt;td&gt;&lt;strong&gt;100%&lt;/strong&gt; (100%)&lt;/td&gt;
					&lt;td&gt;&lt;strong&gt;100%&lt;/strong&gt; (100%)&lt;/td&gt;
					&lt;td&gt;95.8% (100%)&lt;/td&gt;
			&lt;/tr&gt;
			&lt;tr&gt;
					&lt;td&gt;Notification priority&lt;/td&gt;
					&lt;td&gt;95.8% (100%)&lt;/td&gt;
					&lt;td&gt;&lt;strong&gt;100%&lt;/strong&gt; (100%)&lt;/td&gt;
					&lt;td&gt;75.0% (100%)&lt;/td&gt;
			&lt;/tr&gt;
			&lt;tr&gt;
					&lt;td&gt;Tool-call risk&lt;/td&gt;
					&lt;td&gt;&lt;strong&gt;87.5%&lt;/strong&gt; (95.8%)&lt;/td&gt;
					&lt;td&gt;&lt;strong&gt;87.5%&lt;/strong&gt; (95.8%)&lt;/td&gt;
					&lt;td&gt;83.3% (100%)&lt;/td&gt;
			&lt;/tr&gt;
			&lt;tr&gt;
					&lt;td&gt;All five&lt;/td&gt;
					&lt;td&gt;88.6% (93.7%)&lt;/td&gt;
					&lt;td&gt;&lt;strong&gt;92.3%&lt;/strong&gt; (96.6%)&lt;/td&gt;
					&lt;td&gt;82.1% (94.9%)&lt;/td&gt;
			&lt;/tr&gt;
	&lt;/tbody&gt;
&lt;/table&gt;
&lt;p&gt;&lt;em&gt;Accuracy by decision at the production threshold, with accuracy at each model&amp;rsquo;s own best threshold in brackets&lt;/em&gt;&lt;/p&gt;
&lt;figure class="doc-figure"&gt;&lt;img src="https://ankush.one/media/construct/clef-vs-jev-benchmark/02-accuracy-by-decision.webp" alt="Grouped bar chart of accuracy for five decisions. Entity match: Jev 72%, Clef 87%, Clef-flash 80%. Email triage: 93%, 87%, 73%. Memory extract gate: 100%, 100%, 96%. Notification priority: 96%, 100%, 75%. Tool-call risk: 88%, 88%, 83%." title="Accuracy per decision at today&amp;#39;s thresholds. Jev keeps email triage. Clef takes entity matching and notification priority." width="2000" height="1125" loading="lazy" decoding="async"&gt;&lt;figcaption&gt;Accuracy per decision at today&amp;#39;s thresholds. Jev keeps email triage. Clef takes entity matching and notification priority.&lt;/figcaption&gt;&lt;/figure&gt;&lt;p&gt;The honest size of this result: Clef leads Jev by 3.7 points on 117 cases, with a 95% interval of -1.1 to +8.8 points. That is not statistically significant, and we would not stretch it into &amp;ldquo;Clef beats Jev&amp;rdquo;. Three of the entity cases are ones our product resolves in code before any model is asked (identical names, and names that differ only in case). Jev did worst on exactly those, and without them the lead shrinks to 2.0 points (92.1% against 90.1%). What we can say is narrower and more useful. Cloudflare&amp;rsquo;s public numbers predicted Clef would be at least as good as Jev on decision tasks, and on private prompts it had never seen, written for its competitor, it was not measurably worse.&lt;/p&gt;
&lt;p&gt;Jev deserves its due here too. It won email triage outright, it never made a wrong merge, and it was perfect on the extract gate. This is a close race between two good models.&lt;/p&gt;
&lt;h2 id="is-clef-deterministic"&gt;Is Clef deterministic?&lt;a class="heading-anchor" href="#is-clef-deterministic" aria-label="Link to this section"&gt;&lt;/a&gt;
&lt;/h2&gt;
&lt;p&gt;Yes, on everything we sent it. We called each model three times with the same request for all 117 cases. Clef returned the same probability, to four decimal places, on all 351 calls. So did Clef-flash. To rule out a cache, we sent the same requests again with changes that do not touch the model&amp;rsquo;s input (a unique query string, reformatted JSON, reordered keys): same answers, a cache miss on every response, and a new inference ID each time.&lt;/p&gt;
&lt;figure class="doc-figure"&gt;&lt;img src="https://ankush.one/media/construct/clef-vs-jev-benchmark/05-repeatability.webp" alt="Bar chart of the largest probability change between three identical calls, by decision. Jev: 0.07, 0.04, 0.03, 0.10, 0.04. Clef and Clef-flash: 0.00 on every decision." title="Largest change in probability between three identical calls. Clef and Clef-flash sit at zero on every decision." width="2000" height="1125" loading="lazy" decoding="async"&gt;&lt;figcaption&gt;Largest change in probability between three identical calls. Clef and Clef-flash sit at zero on every decision.&lt;/figcaption&gt;&lt;/figure&gt;&lt;p&gt;The probability each Jev decision acts on moved between identical calls by up to 0.10, and two cases changed verdict from one call to the next. Across every probability Jev returned, the largest move was 0.16, and 100 of the 117 cases came back with at least one number different. TypeSafe describes Jev as returning &amp;ldquo;similar answers for similar inputs&amp;rdquo;, which is a promise of consistency, not determinism, and a wobble of this size is consistent. But a decision layer acts when a probability crosses a line, so any wobble means a case near the line is decided by which call you happened to make.&lt;/p&gt;
&lt;p&gt;Cloudflare&amp;rsquo;s post explains why Clef behaves differently: the decision step is non-autoregressive, with choices derived directly from the model&amp;rsquo;s internal representations and no text sampled along the way. We can confirm the practical result. For anyone who has debugged an agent that did something different the second time, this is the finding that matters most, more than a few points of accuracy.&lt;/p&gt;
&lt;h2 id="jevs-scores-moved-under-the-same-version-name"&gt;Jev&amp;rsquo;s scores moved under the same version name&lt;a class="heading-anchor" href="#jevs-scores-moved-under-the-same-version-name" aria-label="Link to this section"&gt;&lt;/a&gt;
&lt;/h2&gt;
&lt;p&gt;This one surprised us. On September 20 we calibrated our live entity matcher on Jev. Real matches scored 0.71 to 0.82, everything else scored 0.34 or less, and we put the merge threshold at 0.70, in the gap.&lt;/p&gt;
&lt;p&gt;Two weeks later we asked the same three calibration probes again.&lt;/p&gt;
&lt;figure class="doc-figure"&gt;&lt;img src="https://ankush.one/media/construct/clef-vs-jev-benchmark/04-jev-drift.webp" alt="Dot plot of Jev’s probability on three calibration probes. Dr. Priya Shah equals Priya Shah: 0.71 on 20 September; on 4 October, 0.47 through Workers AI and 0.46 on TypeSafe’s own API. Priya Shah equals Priya Shah: 0.76, then 0.42 and 0.43. Atlas migration: 0.82, then 0.60 and 0.61. The merge threshold is 0.70." title="Jev&amp;#39;s probability for the correct match on our three calibration probes: one call on September 20, and the mean of 20 calls on October 4 through Workers AI and on TypeSafe&amp;#39;s own API with the version pinned. All were served as jev-1.13.0." width="2000" height="1125" loading="lazy" decoding="async"&gt;&lt;figcaption&gt;Jev&amp;#39;s probability for the correct match on our three calibration probes: one call on September 20, and the mean of 20 calls on October 4 through Workers AI and on TypeSafe&amp;#39;s own API with the version pinned. All were served as jev-1.13.0.&lt;/figcaption&gt;&lt;/figure&gt;&lt;p&gt;All three now fall under the bar. To be sure this was not a noisy afternoon, we asked each probe 20 more times: none of the 60 calls reached 0.70, and the highest was 0.66. Clef, asked the same 20 times, returned the identical number every time. The served version was &lt;code&gt;jev-1.13.0&lt;/code&gt; both times, so the alarm we built to catch version changes stayed quiet. On this benchmark Jev still never merged two different people, which is the expensive mistake, but it made only 44% of the merges it should have. In production, a missed merge becomes a duplicate.&lt;/p&gt;
&lt;p&gt;We do not know why the scores moved. Our first suspect was our own route: we call Jev through Cloudflare&amp;rsquo;s Workers AI, which cannot pin a version. So we repeated everything against TypeSafe&amp;rsquo;s own API with the version pinned to &lt;code&gt;jev-1.13.0&lt;/code&gt;. Same result: none of 60 calls reached 0.70, and Jev&amp;rsquo;s answers still varied between identical calls on 104 of 117 cases. The route is not the cause. Three probes are still three probes, and our September baseline is one call each. We are not claiming Jev got worse in general. The lesson is about our own setup: a threshold is only as stable as the model under it, and a version string is not a guarantee. Open weights can change that, if you use them: a model you pin or host yourself cannot move under its thresholds. The hosted Clef alias we tested carries no version either, so we will re-ask these same probes on Clef in two weeks.&lt;/p&gt;
&lt;h2 id="entity-matching-where-the-three-models-differ-most"&gt;Entity matching: where the three models differ most&lt;a class="heading-anchor" href="#entity-matching-where-the-three-models-differ-most" aria-label="Link to this section"&gt;&lt;/a&gt;
&lt;/h2&gt;
&lt;p&gt;Entity matching is the decision where a mistake is permanent. A wrong merge fuses two real people in the agent&amp;rsquo;s memory, so precision matters more than recall. &lt;a href="https://construct.computer/blog/ai-agent-memory/" data-external&gt;AI agent memory&lt;/a&gt; covers what that memory holds and how you can correct it.&lt;/p&gt;
&lt;figure class="doc-figure"&gt;&lt;img src="https://ankush.one/media/construct/clef-vs-jev-benchmark/03-entity-match.webp" alt="Bar chart for entity matching at a 0.70 threshold. Correct merges made: Jev 44%, Clef 80%, Clef-flash 80%. Merges that were right: Jev 100%, Clef 92%, Clef-flash 80%." title="Entity matching at the production threshold of 0.70: 15 pairs that should merge and 15 that should not." width="2000" height="1125" loading="lazy" decoding="async"&gt;&lt;figcaption&gt;Entity matching at the production threshold of 0.70: 15 pairs that should merge and 15 that should not.&lt;/figcaption&gt;&lt;/figure&gt;&lt;ul&gt;
&lt;li&gt;&lt;strong&gt;Clef&lt;/strong&gt; caught the aliases this decision exists for: &amp;ldquo;Jose Garcia Marquez&amp;rdquo; against the accented spelling at 0.98, &amp;ldquo;Stripe, Inc.&amp;rdquo; against &amp;ldquo;Stripe&amp;rdquo; at 0.98, &amp;ldquo;I.B.M.&amp;rdquo; against &amp;ldquo;IBM&amp;rdquo; at 0.99. Its one wrong merge was &amp;ldquo;Acme&amp;rdquo; into a lone &amp;ldquo;Acme Corporation&amp;rdquo; at 0.78, a case reasonable people would argue about.&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;At a 0.80 threshold, Clef made 10 of 15 correct merges and no wrong ones.&lt;/strong&gt; That is our starting point for shadow mode, not a recommendation: it was picked on these same 30 cases, and &amp;ldquo;Acme&amp;rdquo; against &amp;ldquo;Acme Corporation&amp;rdquo; sits at 0.78, just under it.&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Jev is not worse at telling people apart.&lt;/strong&gt; Drop its threshold to 0.50 and it makes 64% of the correct merges, still with no wrong ones, close to Clef&amp;rsquo;s 67% at 0.80. Its problem at 0.70 is that its scores moved, not that it lost the ability to rank the pairs.&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Clef-flash&lt;/strong&gt; is quick to say yes. It matched &amp;ldquo;Acme Labs Europe&amp;rdquo; to &amp;ldquo;Acme Labs&amp;rdquo; at 0.76 and the bare concept &amp;ldquo;pricing&amp;rdquo; to &amp;ldquo;pricing model&amp;rdquo; at 0.89. For a merge, that is too eager. For a gate that only skips work, it is fine.&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;All three&lt;/strong&gt; refused a one-letter surname typo, &amp;ldquo;Mendelsohn&amp;rdquo; against &amp;ldquo;Mendelson&amp;rdquo;. Good. Those can be two people.&lt;/li&gt;
&lt;/ul&gt;
&lt;h2 id="thresholds-do-not-transfer-between-decision-models"&gt;Thresholds do not transfer between decision models&lt;a class="heading-anchor" href="#thresholds-do-not-transfer-between-decision-models" aria-label="Link to this section"&gt;&lt;/a&gt;
&lt;/h2&gt;
&lt;p&gt;If you take one thing from this post into your own migration, take this. A threshold is a property of a model and a prompt together. Move the model and the threshold is wrong, even when the new model is better.&lt;/p&gt;
&lt;p&gt;The three models are not far apart in ability. Asked to rank our cases from &amp;ldquo;should act&amp;rdquo; to &amp;ldquo;should not&amp;rdquo;, all three do it well: on the four decisions we could measure this way, the ranking score (AUROC) is 0.93 or higher for every model and a perfect 1.00 on two decisions. What differs is the number each model attaches to the same judgment.&lt;/p&gt;
&lt;table&gt;
	&lt;thead&gt;
			&lt;tr&gt;
					&lt;th&gt;Decision&lt;/th&gt;
					&lt;th&gt;Configured&lt;/th&gt;
					&lt;th&gt;Jev&amp;rsquo;s best&lt;/th&gt;
					&lt;th&gt;Clef&amp;rsquo;s best&lt;/th&gt;
					&lt;th&gt;Clef-flash&amp;rsquo;s best&lt;/th&gt;
			&lt;/tr&gt;
	&lt;/thead&gt;
	&lt;tbody&gt;
			&lt;tr&gt;
					&lt;td&gt;Tool-call risk&lt;/td&gt;
					&lt;td&gt;0.80&lt;/td&gt;
					&lt;td&gt;0.49&lt;/td&gt;
					&lt;td&gt;0.22&lt;/td&gt;
					&lt;td&gt;0.31&lt;/td&gt;
			&lt;/tr&gt;
			&lt;tr&gt;
					&lt;td&gt;Notification priority&lt;/td&gt;
					&lt;td&gt;0.85&lt;/td&gt;
					&lt;td&gt;0.60&lt;/td&gt;
					&lt;td&gt;0.85&lt;/td&gt;
					&lt;td&gt;0.61&lt;/td&gt;
			&lt;/tr&gt;
	&lt;/tbody&gt;
&lt;/table&gt;
&lt;p&gt;&lt;em&gt;The most accurate threshold per model on two decisions, against the configured threshold&lt;/em&gt;&lt;/p&gt;
&lt;p&gt;Clef-flash shows the effect best. On Jev&amp;rsquo;s thresholds it scored 82.1%. With thresholds chosen for it on held-out cases it scored 90.6%, level with Jev&amp;rsquo;s 90.9% on the same footing, and Clef stayed ahead at 93.2%. The 9B model was never bad at these decisions. It was being read with another model&amp;rsquo;s ruler.&lt;/p&gt;
&lt;p&gt;The thresholds in the table were found on the same cases they are scored on, which is why the bracketed figures in the results table (96.6%, 94.9%, 93.7%) run higher than the held-out ones. Treat them as a direction and not as settings. The practical rule is simple: swap the model in shadow, log probabilities for a week, and set the threshold from that.&lt;/p&gt;
&lt;h2 id="how-fast-is-clef-what-we-could-and-could-not-measure"&gt;How fast is Clef? What we could and could not measure&lt;a class="heading-anchor" href="#how-fast-is-clef-what-we-could-and-could-not-measure" aria-label="Link to this section"&gt;&lt;/a&gt;
&lt;/h2&gt;
&lt;p&gt;Cloudflare reports a median of 209 ms for Clef and 38.8 ms for Clef-flash, against 524 ms for Jev. We are not going to confirm or dispute those numbers, because our harness was the wrong instrument.&lt;/p&gt;
&lt;p&gt;We called all three models through Cloudflare&amp;rsquo;s REST API from a laptop. The fastest call we saw for any model was 340 ms, and for every model the floor sat between 340 and 384 ms. That floor is the path: authentication, the public API, the gateway, and the round trip. You cannot see a 38 ms model through a 340 ms window.&lt;/p&gt;
&lt;table&gt;
	&lt;thead&gt;
			&lt;tr&gt;
					&lt;th&gt;&lt;/th&gt;
					&lt;th&gt;Jev 1.13&lt;/th&gt;
					&lt;th&gt;Clef&lt;/th&gt;
					&lt;th&gt;Clef-flash&lt;/th&gt;
			&lt;/tr&gt;
	&lt;/thead&gt;
	&lt;tbody&gt;
			&lt;tr&gt;
					&lt;td&gt;Fastest call&lt;/td&gt;
					&lt;td&gt;384 ms&lt;/td&gt;
					&lt;td&gt;352 ms&lt;/td&gt;
					&lt;td&gt;340 ms&lt;/td&gt;
			&lt;/tr&gt;
			&lt;tr&gt;
					&lt;td&gt;Median&lt;/td&gt;
					&lt;td&gt;454 ms&lt;/td&gt;
					&lt;td&gt;654 ms&lt;/td&gt;
					&lt;td&gt;449 ms&lt;/td&gt;
			&lt;/tr&gt;
			&lt;tr&gt;
					&lt;td&gt;95th percentile&lt;/td&gt;
					&lt;td&gt;591 ms&lt;/td&gt;
					&lt;td&gt;1,262 ms&lt;/td&gt;
					&lt;td&gt;1,024 ms&lt;/td&gt;
			&lt;/tr&gt;
	&lt;/tbody&gt;
&lt;/table&gt;
&lt;p&gt;&lt;em&gt;Latency per call through the REST API from a laptop, 351 calls per model. Not a production measurement.&lt;/em&gt;&lt;/p&gt;
&lt;p&gt;For transparency, those are the raw numbers. On this path Clef-flash matched Jev at the median, Clef was slower, and both Clef models had longer tails, three days after launch. None of that tells you what a Worker sees when it calls the model through the AI binding, next to the GPU, which is how we call it in production. That measurement is next, and we will publish it here.&lt;/p&gt;
&lt;h2 id="what-does-clef-cost-compared-with-jev"&gt;What does Clef cost compared with Jev?&lt;a class="heading-anchor" href="#what-does-clef-cost-compared-with-jev" aria-label="Link to this section"&gt;&lt;/a&gt;
&lt;/h2&gt;
&lt;figure class="doc-figure"&gt;&lt;img src="https://ankush.one/media/construct/clef-vs-jev-benchmark/07-cost.webp" alt="Bar chart of cost per million decisions: Jev 23 dollars at 541 input tokens per call, Clef 93 dollars at 387 tokens, Clef-flash 35 dollars at 387 tokens." title="Cost per million decisions at list price, using the mean input tokens each model counted for our requests." width="2000" height="1125" loading="lazy" decoding="async"&gt;&lt;figcaption&gt;Cost per million decisions at list price, using the mean input tokens each model counted for our requests.&lt;/figcaption&gt;&lt;/figure&gt;&lt;p&gt;Clef&amp;rsquo;s list price is 5.7 times Jev&amp;rsquo;s, which is the first thing the Hacker News thread did arithmetic on. Two things soften that.&lt;/p&gt;
&lt;p&gt;First, Clef counted 28% fewer tokens than Jev for the same request: 387 per call against 541. Per decision, the gap is about 4 times, not 5.7.&lt;/p&gt;
&lt;p&gt;Second, look at the unit. A million Clef decisions cost $93. All 1,053 calls in this benchmark cost about five cents. An agent that makes a thousand decisions a day spends about nine cents a day on Clef and two on Jev, next to a language model bill many times that. If you run billions of decisions a day, price it carefully, and look at Clef-flash at $35 per million. For everyone else, accuracy and repeatability are worth more than the difference.&lt;/p&gt;
&lt;h2 id="clef-vs-jev-which-decision-model-should-you-use"&gt;Clef vs Jev: which decision model should you use?&lt;a class="heading-anchor" href="#clef-vs-jev-which-decision-model-should-you-use" aria-label="Link to this section"&gt;&lt;/a&gt;
&lt;/h2&gt;
&lt;table&gt;
	&lt;thead&gt;
			&lt;tr&gt;
					&lt;th&gt;If you need&lt;/th&gt;
					&lt;th&gt;Pick&lt;/th&gt;
					&lt;th&gt;Why&lt;/th&gt;
			&lt;/tr&gt;
	&lt;/thead&gt;
	&lt;tbody&gt;
			&lt;tr&gt;
					&lt;td&gt;The best accuracy without retuning&lt;/td&gt;
					&lt;td&gt;Clef&lt;/td&gt;
					&lt;td&gt;Highest drop-in accuracy on our decisions, 92.3%&lt;/td&gt;
			&lt;/tr&gt;
			&lt;tr&gt;
					&lt;td&gt;The same answer every time&lt;/td&gt;
					&lt;td&gt;Clef or Clef-flash&lt;/td&gt;
					&lt;td&gt;Identical probabilities on every repeated call&lt;/td&gt;
			&lt;/tr&gt;
			&lt;tr&gt;
					&lt;td&gt;To pin a version, self-host, or fine-tune&lt;/td&gt;
					&lt;td&gt;Clef or Clef-flash&lt;/td&gt;
					&lt;td&gt;Open weights under Apache 2.0&lt;/td&gt;
			&lt;/tr&gt;
			&lt;tr&gt;
					&lt;td&gt;To classify images&lt;/td&gt;
					&lt;td&gt;Clef or Clef-flash&lt;/td&gt;
					&lt;td&gt;Both have a vision encoder; Jev is text only&lt;/td&gt;
			&lt;/tr&gt;
			&lt;tr&gt;
					&lt;td&gt;The lowest price per token&lt;/td&gt;
					&lt;td&gt;Jev&lt;/td&gt;
					&lt;td&gt;$0.042 per million input tokens&lt;/td&gt;
			&lt;/tr&gt;
			&lt;tr&gt;
					&lt;td&gt;Zero wrong merges at your current thresholds&lt;/td&gt;
					&lt;td&gt;Jev&lt;/td&gt;
					&lt;td&gt;100% precision on entity matching, at the cost of recall&lt;/td&gt;
			&lt;/tr&gt;
			&lt;tr&gt;
					&lt;td&gt;A high-volume gate that only skips work&lt;/td&gt;
					&lt;td&gt;Clef-flash, with its own thresholds&lt;/td&gt;
					&lt;td&gt;90.6% once retuned on held-out cases, at about a third of Clef&amp;rsquo;s price&lt;/td&gt;
			&lt;/tr&gt;
			&lt;tr&gt;
					&lt;td&gt;Anything that merges, deletes or sends&lt;/td&gt;
					&lt;td&gt;Clef, at a high threshold&lt;/td&gt;
					&lt;td&gt;Clef-flash was too eager on lookalike names&lt;/td&gt;
			&lt;/tr&gt;
	&lt;/tbody&gt;
&lt;/table&gt;
&lt;p&gt;&lt;em&gt;Which decision model fits which job, based on our benchmark and each vendor&amp;rsquo;s documentation&lt;/em&gt;&lt;/p&gt;
&lt;p&gt;The broader field is moving fast. Besides Jev and Clef there are open decision models such as Laya, Kev and OpenDecider, and Cloudflare&amp;rsquo;s own table includes several of them. We have only tested the three in this post.&lt;/p&gt;
&lt;h2 id="what-we-are-switching-and-why-cloudflare-makes-it-easy"&gt;What we are switching, and why Cloudflare makes it easy&lt;a class="heading-anchor" href="#what-we-are-switching-and-why-cloudflare-makes-it-easy" aria-label="Link to this section"&gt;&lt;/a&gt;
&lt;/h2&gt;
&lt;p&gt;Construct already runs on Cloudflare: Workers, Durable Objects, AI Gateway, and Workers AI for the small models. &lt;a href="https://construct.computer/blog/running-ai-agents-on-cloudflare-not-vms/" data-external&gt;How our agents get computers we mostly do not pay for&lt;/a&gt; describes that stack. So Clef is not a new vendor for us. It is a new model ID on a binding we already call.&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;&lt;strong&gt;Entity matching goes to Clef in shadow first, starting at 0.80.&lt;/strong&gt; Jev keeps deciding while Clef&amp;rsquo;s answers are recorded beside it. The threshold we ship will come from that data, and we switch only if real traffic agrees with this benchmark.&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Background decisions default to Clef,&lt;/strong&gt; each with its own threshold measured in shadow.&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Interactive decisions wait for the latency test.&lt;/strong&gt; If Clef-flash is as fast from a Worker as Cloudflare reports, it is the natural fit for anything a person is waiting on.&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Clef-flash stays away from anything that merges or deletes&lt;/strong&gt; until it has its own thresholds.&lt;/li&gt;
&lt;/ul&gt;
&lt;p&gt;There is a quieter benefit. In our Jev post we said that decisions which read your conversations or email &amp;ldquo;wait for the paperwork&amp;rdquo;, because sending that text to another provider is a privacy decision. A first-party Cloudflare model keeps that text on the platform that already runs the rest of Construct, which removes the main reason most of our decisions are still switched off.&lt;/p&gt;
&lt;p&gt;We are also watching Cloudflare&amp;rsquo;s new fine-tuning work, which it calls Reinforcement Learning for Calibrated Decisions. Our decision layer already records the label, the probability and what the product did for every decision, with no user text. That is close to the dataset such a system wants.&lt;/p&gt;
&lt;p&gt;&lt;em&gt;Read next: &lt;a href="https://construct.computer/blog/jev-ai-agents/" data-external&gt;We put Jev inside our AI employee&lt;/a&gt;&lt;/em&gt;&lt;/p&gt;
&lt;h2 id="how-to-try-clef"&gt;How to try Clef&lt;a class="heading-anchor" href="#how-to-try-clef" aria-label="Link to this section"&gt;&lt;/a&gt;
&lt;/h2&gt;
&lt;ul&gt;
&lt;li&gt;&lt;strong&gt;On Workers AI:&lt;/strong&gt; call &lt;code&gt;@cf/cloudflare/clef&lt;/code&gt; or &lt;code&gt;@cf/cloudflare/clef-flash&lt;/code&gt; through the AI binding or the REST API. The request is a &lt;code&gt;state&lt;/code&gt; plus a map of &lt;code&gt;questions&lt;/code&gt;, the same as Jev.&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Coming from Jev:&lt;/strong&gt; change the model ID and keep everything else. Then re-measure every threshold.&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Self-hosted:&lt;/strong&gt; the weights are at &lt;code&gt;Cloudflare/clef&lt;/code&gt; and &lt;code&gt;Cloudflare/clef-flash&lt;/code&gt; on Hugging Face under Apache 2.0.&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Read the probabilities, not the &lt;code&gt;confidence&lt;/code&gt; field.&lt;/strong&gt; It means different things on Clef and Jev, and on Jev it changed meaning between our September and October runs. Read &lt;code&gt;probabilities&lt;/code&gt; and both models behave the same.&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Pace a REST benchmark.&lt;/strong&gt; Cloudflare&amp;rsquo;s REST API throttled our token at roughly 90 calls a minute. The Workers binding is the route for real traffic.&lt;/li&gt;
&lt;/ul&gt;
&lt;p&gt;A minimal request:&lt;/p&gt;
&lt;div class="highlight"&gt;&lt;pre tabindex="0" class="chroma"&gt;&lt;code class="language-bash" data-lang="bash"&gt;&lt;span class="line"&gt;&lt;span class="cl"&gt;curl https://api.cloudflare.com/client/v4/accounts/&lt;span class="nv"&gt;$ACCOUNT_ID&lt;/span&gt;/ai/run &lt;span class="se"&gt;\
&lt;/span&gt;&lt;/span&gt;&lt;/span&gt;&lt;span class="line"&gt;&lt;span class="cl"&gt; -H &lt;span class="s2"&gt;&amp;#34;Authorization: Bearer &lt;/span&gt;&lt;span class="nv"&gt;$API_TOKEN&lt;/span&gt;&lt;span class="s2"&gt;&amp;#34;&lt;/span&gt; &lt;span class="se"&gt;\
&lt;/span&gt;&lt;/span&gt;&lt;/span&gt;&lt;span class="line"&gt;&lt;span class="cl"&gt; -d &lt;span class="s1"&gt;&amp;#39;{
&lt;/span&gt;&lt;/span&gt;&lt;/span&gt;&lt;span class="line"&gt;&lt;span class="cl"&gt;&lt;span class="s1"&gt; &amp;#34;model&amp;#34;: &amp;#34;@cf/cloudflare/clef&amp;#34;,
&lt;/span&gt;&lt;/span&gt;&lt;/span&gt;&lt;span class="line"&gt;&lt;span class="cl"&gt;&lt;span class="s1"&gt; &amp;#34;input&amp;#34;: {
&lt;/span&gt;&lt;/span&gt;&lt;/span&gt;&lt;span class="line"&gt;&lt;span class="cl"&gt;&lt;span class="s1"&gt; &amp;#34;state&amp;#34;: { &amp;#34;entity&amp;#34;: { &amp;#34;type&amp;#34;: &amp;#34;person&amp;#34;, &amp;#34;name&amp;#34;: &amp;#34;Dr. Priya Shah&amp;#34; } },
&lt;/span&gt;&lt;/span&gt;&lt;/span&gt;&lt;span class="line"&gt;&lt;span class="cl"&gt;&lt;span class="s1"&gt; &amp;#34;questions&amp;#34;: {
&lt;/span&gt;&lt;/span&gt;&lt;/span&gt;&lt;span class="line"&gt;&lt;span class="cl"&gt;&lt;span class="s1"&gt; &amp;#34;match&amp;#34;: {
&lt;/span&gt;&lt;/span&gt;&lt;/span&gt;&lt;span class="line"&gt;&lt;span class="cl"&gt;&lt;span class="s1"&gt; &amp;#34;type&amp;#34;: &amp;#34;choice&amp;#34;,
&lt;/span&gt;&lt;/span&gt;&lt;/span&gt;&lt;span class="line"&gt;&lt;span class="cl"&gt;&lt;span class="s1"&gt; &amp;#34;instructions&amp;#34;: &amp;#34;Which candidate is the same entity?&amp;#34;,
&lt;/span&gt;&lt;/span&gt;&lt;/span&gt;&lt;span class="line"&gt;&lt;span class="cl"&gt;&lt;span class="s1"&gt; &amp;#34;criteria&amp;#34;: {
&lt;/span&gt;&lt;/span&gt;&lt;/span&gt;&lt;span class="line"&gt;&lt;span class="cl"&gt;&lt;span class="s1"&gt; &amp;#34;node-priya&amp;#34;: &amp;#34;Priya Shah (person)&amp;#34;,
&lt;/span&gt;&lt;/span&gt;&lt;/span&gt;&lt;span class="line"&gt;&lt;span class="cl"&gt;&lt;span class="s1"&gt; &amp;#34;none&amp;#34;: &amp;#34;No candidate is clearly the same entity&amp;#34;
&lt;/span&gt;&lt;/span&gt;&lt;/span&gt;&lt;span class="line"&gt;&lt;span class="cl"&gt;&lt;span class="s1"&gt; }
&lt;/span&gt;&lt;/span&gt;&lt;/span&gt;&lt;span class="line"&gt;&lt;span class="cl"&gt;&lt;span class="s1"&gt; }
&lt;/span&gt;&lt;/span&gt;&lt;/span&gt;&lt;span class="line"&gt;&lt;span class="cl"&gt;&lt;span class="s1"&gt; }
&lt;/span&gt;&lt;/span&gt;&lt;/span&gt;&lt;span class="line"&gt;&lt;span class="cl"&gt;&lt;span class="s1"&gt; }
&lt;/span&gt;&lt;/span&gt;&lt;/span&gt;&lt;span class="line"&gt;&lt;span class="cl"&gt;&lt;span class="s1"&gt; }&amp;#39;&lt;/span&gt;
&lt;/span&gt;&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;&lt;/div&gt;&lt;h2 id="what-this-benchmark-does-not-show"&gt;What this benchmark does not show&lt;a class="heading-anchor" href="#what-this-benchmark-does-not-show" aria-label="Link to this section"&gt;&lt;/a&gt;
&lt;/h2&gt;
&lt;ul&gt;
&lt;li&gt;&lt;strong&gt;It is small.&lt;/strong&gt; 117 cases, 15 to 30 per decision. A 3.7 point lead at that size is suggestive.&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;It is synthetic.&lt;/strong&gt; The cases are written to cover each decision&amp;rsquo;s edges. Production traffic is mostly the easy middle, so real accuracy for all three models is probably higher.&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;The labels are ours, drafted with an AI assistant.&lt;/strong&gt; No second person has checked them yet. A few are judgment calls, like whether &amp;ldquo;Acme&amp;rdquo; should merge into the only &amp;ldquo;Acme Corporation&amp;rdquo; in memory. We said no. Dropping the five most arguable ones does not change the order.&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Thresholds are tuned on small sets.&lt;/strong&gt; The held-out figures are the fair ones, and they are still noisy at 15 to 30 cases per decision.&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Latency is unmeasured.&lt;/strong&gt; See above.&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;One day.&lt;/strong&gt; Everything ran on October 4, through Cloudflare, with Jev repeated on TypeSafe&amp;rsquo;s own API. Clef was three days old. Jev&amp;rsquo;s scores moved in two weeks. Either could look different next month, which is the argument for running your own cases on a schedule.&lt;/li&gt;
&lt;/ul&gt;
&lt;p&gt;The paper has the statistics behind each of these, the five cases no model got right, and every case where any model slipped. The cases, labels and every call are &lt;a href="https://construct.computer/blog/clef-vs-jev-benchmark/data/cases.json" data-external&gt;published as JSON&lt;/a&gt; beside it: &lt;a href="https://construct.computer/blog/clef-vs-jev-benchmark/compatible-apis-incompatible-thresholds.pdf" data-external&gt;Compatible APIs, Incompatible Thresholds: A Drop-In Replacement Study of Clef and Jev on Decision Specifications from a Production Agent&lt;/a&gt;.&lt;/p&gt;
&lt;p&gt;Cloudflare published benchmarks, and people on the internet said benchmarks are easy to cheat. The fix for that is boring: write down your own cases, label them, and run every new model through them. Ours took an afternoon and a few cents, and it told us more about our own system than any leaderboard could.&lt;/p&gt;
&lt;h2 id="where-construct-fits"&gt;Where Construct fits&lt;a class="heading-anchor" href="#where-construct-fits" aria-label="Link to this section"&gt;&lt;/a&gt;
&lt;/h2&gt;
&lt;p&gt;Construct is an AI employee with its own computer: a workspace filesystem, long-term memory, a schedule, email, and workflows you can read before you run them. Decision models make the small calls around the edges, such as whether a person your agent just read about is one it already knows. Planning and writing run on general-purpose language models.&lt;/p&gt;
&lt;p&gt;Two boundaries, stated plainly. As of this post, Jev still makes the live entity-matching decision and Clef has only run in this benchmark; this page will say when that changes. And the other four decisions in this benchmark are built but switched off in production until shadow data supports them.&lt;/p&gt;
&lt;p&gt;If decision models are new to you, start with &lt;a href="https://construct.computer/blog/jev-ai-agents/" data-external&gt;we put Jev inside our AI employee&lt;/a&gt;. For why small, repeatable judgments matter so much in long tasks, read &lt;a href="https://construct.computer/blog/agent-task-half-life/" data-external&gt;your agent has a half-life&lt;/a&gt;. And if the category itself is new, begin with &lt;a href="https://construct.computer/blog/what-is-an-ai-employee/" data-external&gt;what is an AI employee&lt;/a&gt;.&lt;/p&gt;
&lt;hr&gt;
&lt;p&gt;&lt;em&gt;Originally posted on &lt;a href="https://construct.computer/blog/clef-vs-jev-benchmark/" data-external&gt;construct.computer&lt;/a&gt;.&lt;/em&gt;&lt;/p&gt;</content:encoded></item><item><title>We put Jev inside our AI employee</title><link>https://ankush.one/blogs/jev-ai-agents/</link><pubDate>Wed, 30 Sep 2026 00:00:00 +0000</pubDate><author>ankush4singh@gmail.com (Ankush Singh)</author><guid>https://ankush.one/blogs/jev-ai-agents/</guid><category>ai-agents</category><category>machine-learning</category><category>construct</category><description>We shipped TypeSafe&amp;rsquo;s Jev into our AI employee&amp;rsquo;s memory five days after launch: measured latency, how we set a 0.7 threshold, where Jev fails, and what it costs.</description><content:encoded>&lt;p&gt;Your AI employee reads an email that mentions &amp;ldquo;Dr. Priya Shah&amp;rdquo;. Last week it met &amp;ldquo;Priya Shah&amp;rdquo; in a calendar invite. Same person? Guess yes and get it wrong, and two people&amp;rsquo;s histories are welded together in its memory. Guess no and get it wrong, and you have a duplicate. It makes that call every time it learns something, in the background, where nobody is watching.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Jev is a System One model from TypeSafe AI: you give it a state and a set of typed questions, and instead of generating text it returns a probability for each possible answer, in well under a second, for $0.042 per million input tokens with output free.&lt;/strong&gt; TypeSafe released it in early access on September 15, 2026. Five days later we shipped it into Construct&amp;rsquo;s memory to make exactly the call above. This post covers what we measured, how we picked the threshold, what Jev gets wrong, what it costs next to a language model, and where a decision model fits in an AI agent&amp;rsquo;s loop.&lt;/p&gt;
&lt;p&gt;We build Construct, an AI employee with its own computer. The numbers in this post come from our own integration, measured on September 20, 2026 against &lt;code&gt;jev-1.13.0&lt;/code&gt;. Everything else about Jev comes from TypeSafe&amp;rsquo;s launch post and docs and from independent tests, all checked on &lt;strong&gt;September 30, 2026&lt;/strong&gt;.&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;&lt;strong&gt;Latency:&lt;/strong&gt; 358 to 611 ms for one question on about 400 tokens, and 377 to 544 ms for seven questions on the same state.&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Consistency:&lt;/strong&gt; repeating the same call moved Jev&amp;rsquo;s probabilities by 0.02 at most.&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Threshold:&lt;/strong&gt; we merge at 0.7. In our probe set, real matches scored 0.71 to 0.82 and everything else 0.34 or less.&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Calls:&lt;/strong&gt; one Jev call replaced up to three sequential calls to a generative judge for each entity.&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Cost:&lt;/strong&gt; about $0.017 per 1,000 decisions of 400 input tokens at list price, with output free.&lt;/li&gt;
&lt;/ul&gt;
&lt;h2 id="what-is-jev"&gt;What is Jev?&lt;a class="heading-anchor" href="#what-is-jev" aria-label="Link to this section"&gt;&lt;/a&gt;
&lt;/h2&gt;
&lt;p&gt;Jev is a decision model, not a chat model: it never generates text, and TypeSafe says it is not a large language model. TypeSafe AI, a San Francisco company founded by Diogo Almeida (a former OpenAI researcher and co-author of the InstructGPT paper), Erik Gafni, and Sasha Sheng, released it alongside a $40 million seed round led by DCVC. TypeSafe calls Jev the first &amp;ldquo;System One model&amp;rdquo;, after the fast, intuitive System 1 in Daniel Kahneman&amp;rsquo;s &lt;em&gt;Thinking, Fast and Slow&lt;/em&gt;, and describes it as &amp;ldquo;a frontier-intelligence function call: unstructured state in, typed probabilistic decisions out&amp;rdquo; (&lt;a href="https://typesafe.ai/blog/introducing-system-one-models-and-jev" data-external&gt;TypeSafe&lt;/a&gt;).&lt;/p&gt;
&lt;p&gt;You send one &lt;code&gt;state&lt;/code&gt;, as text or JSON, and a map of named questions. Each question has one of three types (&lt;a href="https://docs.typesafe.ai/api" data-external&gt;TypeSafe API reference&lt;/a&gt;):&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;&lt;strong&gt;Choice&lt;/strong&gt; picks one of up to 255 labels you define and returns a probability for every label.&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Score&lt;/strong&gt; rates the state on an ordered scale of 2 to 10 levels.&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Noul&lt;/strong&gt;, short for Bernoulli, returns the probability that a statement is true.&lt;/li&gt;
&lt;/ul&gt;
&lt;p&gt;Every question is answered against the same state in one parallel pass, so adding questions barely changes the response time. Jev cannot write a sentence, and that is the point: there is no prose to parse, no output tokens to pay for, and no way for an answer to fall outside the labels you gave it.&lt;/p&gt;
&lt;table&gt;
	&lt;thead&gt;
			&lt;tr&gt;
					&lt;th&gt;&lt;/th&gt;
					&lt;th&gt;Jev (System One model)&lt;/th&gt;
					&lt;th&gt;Language model (LLM)&lt;/th&gt;
			&lt;/tr&gt;
	&lt;/thead&gt;
	&lt;tbody&gt;
			&lt;tr&gt;
					&lt;td&gt;What comes back&lt;/td&gt;
					&lt;td&gt;A typed answer with a probability for each option&lt;/td&gt;
					&lt;td&gt;Generated text, optionally JSON&lt;/td&gt;
			&lt;/tr&gt;
			&lt;tr&gt;
					&lt;td&gt;Can it write, plan, or call tools&lt;/td&gt;
					&lt;td&gt;No&lt;/td&gt;
					&lt;td&gt;Yes&lt;/td&gt;
			&lt;/tr&gt;
			&lt;tr&gt;
					&lt;td&gt;Latency&lt;/td&gt;
					&lt;td&gt;70 to 500 ms claimed; 358 to 611 ms in our tests&lt;/td&gt;
					&lt;td&gt;Usually seconds for a short structured answer&lt;/td&gt;
			&lt;/tr&gt;
			&lt;tr&gt;
					&lt;td&gt;Price per million input tokens&lt;/td&gt;
					&lt;td&gt;$0.042, output free&lt;/td&gt;
					&lt;td&gt;From about $0.05 to $10, plus billed output&lt;/td&gt;
			&lt;/tr&gt;
			&lt;tr&gt;
					&lt;td&gt;Answer outside your options&lt;/td&gt;
					&lt;td&gt;Impossible by construction&lt;/td&gt;
					&lt;td&gt;Possible without a strict schema&lt;/td&gt;
			&lt;/tr&gt;
			&lt;tr&gt;
					&lt;td&gt;Wrong answer&lt;/td&gt;
					&lt;td&gt;Possible, with a probability attached&lt;/td&gt;
					&lt;td&gt;Possible, usually with no probability&lt;/td&gt;
			&lt;/tr&gt;
			&lt;tr&gt;
					&lt;td&gt;Input&lt;/td&gt;
					&lt;td&gt;Text and JSON, 64k tokens per request&lt;/td&gt;
					&lt;td&gt;Text, and often images and files&lt;/td&gt;
			&lt;/tr&gt;
	&lt;/tbody&gt;
&lt;/table&gt;
&lt;p&gt;&lt;em&gt;Jev compared with a general-purpose language model&lt;/em&gt;&lt;/p&gt;
&lt;p&gt;In Kahneman&amp;rsquo;s terms, the language model is the slow, deliberate System 2 and Jev is the fast System 1. The analogy used to run the other way: in 2023, plain LLMs were the System 1 of the story. Simon Willison and Maggie Appleton suggest the plainer name &amp;ldquo;decision models&amp;rdquo; (&lt;a href="https://simonwillison.net/2026/Sep/21/jev/" data-external&gt;Simon Willison&lt;/a&gt;), which is also the term our own code uses.&lt;/p&gt;
&lt;h2 id="using-jev-for-ai-agent-memory"&gt;Using Jev for AI agent memory&lt;a class="heading-anchor" href="#using-jev-for-ai-agent-memory" aria-label="Link to this section"&gt;&lt;/a&gt;
&lt;/h2&gt;
&lt;p&gt;Construct&amp;rsquo;s agent keeps a long-term memory: a graph of the people, companies, and projects in your work, with facts attached to each. After a conversation, it extracts what is worth keeping, and for every person or project it mentions it has to decide whether that is someone it already knows. That is entity resolution, and it is a textbook System One question: a judgment a knowledgeable person makes in a second, asked over and over, in the background, with nobody waiting on the answer. We are not the only ones who ended up here: a September paper, Jev-Mem, proposes putting a System One model in control of an agent&amp;rsquo;s memory (&lt;a href="https://arxiv.org/abs/2609.23986" data-external&gt;arXiv:2609.23986&lt;/a&gt;).&lt;/p&gt;
&lt;p&gt;Most cases never need a model. A platform ID, an exact alias, or a single exact-name match resolves in code. What is left is the ambiguous part, where several candidates share a name or the only matches are fuzzy or come from vector search.&lt;/p&gt;
&lt;p&gt;Before Jev, those leftovers went to a generative judge: a small open model, Llama 4 Scout on Workers AI, constrained by a JSON schema whose only allowed values were the candidate IDs or null. It was asked about each candidate set in turn, up to three sequential calls per entity and up to 64 entities per batch, and it answered with a label and nothing else.&lt;/p&gt;
&lt;p&gt;Now one Jev call scores every candidate at once, whichever lookup found it, plus an explicit &amp;ldquo;none&amp;rdquo;, and entities resolve in parallel, eight at a time, so a 64-entity batch does not open 64 connections at once and run into TypeSafe&amp;rsquo;s per-account rate limit. If Jev does not answer within two seconds, or answers in a shape we did not ask for, resolution falls back to the old judge.&lt;/p&gt;
&lt;p&gt;Who decides what matters, so here it is plainly:&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;&lt;strong&gt;Code&lt;/strong&gt; finds the candidates with vector, exact-name, and fuzzy-name lookups, enforces the threshold, and performs the merge.&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Jev&lt;/strong&gt; answers one question: which candidate, if any, is the same real entity, with a probability for each.&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;A language model&lt;/strong&gt; extracted the entity from the conversation earlier. It takes no part in the decision.&lt;/li&gt;
&lt;/ul&gt;
&lt;p&gt;That split is also our answer to the fair criticism that a decision model has no context. It does not need the whole conversation. Retrieval builds a small state, and Jev makes one bounded call on it. TypeSafe&amp;rsquo;s CEO has said &amp;ldquo;the hard part for coding is actually state engineering&amp;rdquo; (&lt;a href="https://news.ycombinator.com/item?id=49718867" data-external&gt;Hacker News&lt;/a&gt;). For an agent with its own computer and memory, most of that state already exists, structured, before the question is asked.&lt;/p&gt;
&lt;h2 id="how-fast-is-jev-in-production"&gt;How fast is Jev in production?&lt;a class="heading-anchor" href="#how-fast-is-jev-in-production" aria-label="Link to this section"&gt;&lt;/a&gt;
&lt;/h2&gt;
&lt;table&gt;
	&lt;thead&gt;
			&lt;tr&gt;
					&lt;th&gt;Call&lt;/th&gt;
					&lt;th&gt;Input&lt;/th&gt;
					&lt;th&gt;Latency&lt;/th&gt;
			&lt;/tr&gt;
	&lt;/thead&gt;
	&lt;tbody&gt;
			&lt;tr&gt;
					&lt;td&gt;One question&lt;/td&gt;
					&lt;td&gt;About 400 tokens&lt;/td&gt;
					&lt;td&gt;358 to 611 ms&lt;/td&gt;
			&lt;/tr&gt;
			&lt;tr&gt;
					&lt;td&gt;Seven questions on the same state&lt;/td&gt;
					&lt;td&gt;About 711 tokens&lt;/td&gt;
					&lt;td&gt;377 to 544 ms&lt;/td&gt;
			&lt;/tr&gt;
	&lt;/tbody&gt;
&lt;/table&gt;
&lt;p&gt;&lt;em&gt;Jev latency measured by Construct on September 20, 2026&lt;/em&gt;&lt;/p&gt;
&lt;p&gt;Both rows were measured against &lt;code&gt;jev-1.13.0&lt;/code&gt; through Cloudflare AI Gateway. Repeating the seven-question call moved the probabilities by 0.02 at most, and p95 latency was about 0.6 seconds, so we gave the call a two-second budget before falling back. TypeSafe says its service runs on the US West Coast, so your numbers depend on where you call from: OpenRouter measured a 175 ms median (&lt;a href="https://openrouter.ai/blog/insights/jev-vs-claude-opus-5-classification/" data-external&gt;OpenRouter&lt;/a&gt;), and a tester in Japan measured a 245 ms median (&lt;a href="https://towardsdatascience.com/jev-vs-llms-when-ai-moves-from-generation-to-decision-making/" data-external&gt;Towards Data Science&lt;/a&gt;).&lt;/p&gt;
&lt;p&gt;The second row is the useful one. Seven questions took no longer than one and used 1.8 times the input tokens, because Jev reads the state once and answers every question in parallel. Batch every judgment you have about the same state into one call instead of making several. TypeSafe&amp;rsquo;s own cookbook reports 13 batched questions running 12.2 times cheaper and 10 times faster than separate calls, with no change in answers (&lt;a href="https://docs.typesafe.ai/" data-external&gt;TypeSafe docs&lt;/a&gt;).&lt;/p&gt;
&lt;h2 id="how-we-picked-the-07-threshold"&gt;How we picked the 0.7 threshold&lt;a class="heading-anchor" href="#how-we-picked-the-07-threshold" aria-label="Link to this section"&gt;&lt;/a&gt;
&lt;/h2&gt;
&lt;p&gt;Jev returns a probability for every candidate and for &amp;ldquo;none&amp;rdquo;. The real engineering question is where to draw the line. These are the probes we measured on September 20:&lt;/p&gt;
&lt;table&gt;
	&lt;thead&gt;
			&lt;tr&gt;
					&lt;th&gt;New mention&lt;/th&gt;
					&lt;th&gt;Existing candidates&lt;/th&gt;
					&lt;th&gt;Jev&amp;rsquo;s answer&lt;/th&gt;
					&lt;th&gt;Outcome at 0.7&lt;/th&gt;
			&lt;/tr&gt;
	&lt;/thead&gt;
	&lt;tbody&gt;
			&lt;tr&gt;
					&lt;td&gt;&amp;ldquo;Dr. Priya Shah&amp;rdquo;&lt;/td&gt;
					&lt;td&gt;Priya Shah&lt;/td&gt;
					&lt;td&gt;Priya Shah, 0.71&lt;/td&gt;
					&lt;td&gt;Merged&lt;/td&gt;
			&lt;/tr&gt;
			&lt;tr&gt;
					&lt;td&gt;&amp;ldquo;Priya Shah&amp;rdquo;&lt;/td&gt;
					&lt;td&gt;Priya Shah&lt;/td&gt;
					&lt;td&gt;Priya Shah, 0.76&lt;/td&gt;
					&lt;td&gt;Merged&lt;/td&gt;
			&lt;/tr&gt;
			&lt;tr&gt;
					&lt;td&gt;&amp;ldquo;atlas migration&amp;rdquo;&lt;/td&gt;
					&lt;td&gt;Atlas Migration&lt;/td&gt;
					&lt;td&gt;Atlas Migration, 0.82&lt;/td&gt;
					&lt;td&gt;Merged&lt;/td&gt;
			&lt;/tr&gt;
			&lt;tr&gt;
					&lt;td&gt;&amp;ldquo;Acme&amp;rdquo;&lt;/td&gt;
					&lt;td&gt;Acme Corporation&lt;/td&gt;
					&lt;td&gt;None, 0.66 (Acme Corporation 0.34)&lt;/td&gt;
					&lt;td&gt;Kept separate&lt;/td&gt;
			&lt;/tr&gt;
			&lt;tr&gt;
					&lt;td&gt;&amp;ldquo;Acme&amp;rdquo;&lt;/td&gt;
					&lt;td&gt;Acme Corporation, Acme Labs&lt;/td&gt;
					&lt;td&gt;None, 0.79&lt;/td&gt;
					&lt;td&gt;Kept separate&lt;/td&gt;
			&lt;/tr&gt;
			&lt;tr&gt;
					&lt;td&gt;&amp;ldquo;Sam&amp;rdquo;&lt;/td&gt;
					&lt;td&gt;Sam Altman, Samantha Reyes&lt;/td&gt;
					&lt;td&gt;None, 1.00&lt;/td&gt;
					&lt;td&gt;Kept separate&lt;/td&gt;
			&lt;/tr&gt;
			&lt;tr&gt;
					&lt;td&gt;&amp;ldquo;Marcus Webb&amp;rdquo;&lt;/td&gt;
					&lt;td&gt;Unrelated people&lt;/td&gt;
					&lt;td&gt;None, 1.00&lt;/td&gt;
					&lt;td&gt;Kept separate&lt;/td&gt;
			&lt;/tr&gt;
	&lt;/tbody&gt;
&lt;/table&gt;
&lt;p&gt;&lt;em&gt;Jev&amp;rsquo;s answers on entity-resolution probes, and the outcome at a 0.7 merge threshold&lt;/em&gt;&lt;/p&gt;
&lt;p&gt;In this probe set, real matches landed between 0.71 and 0.82, and everything else put 0.34 or less on its best candidate. We set the merge threshold at 0.7, inside that gap.&lt;/p&gt;
&lt;p&gt;Two things decided where in the gap. First, the mistakes are not symmetrical. A wrong merge fuses two real entities and cannot be split again, while a missed merge leaves a duplicate that splits one entity&amp;rsquo;s facts across two records. Second, raising the bar to 0.8 would have refused exactly the alias cases, like &amp;ldquo;Dr. Priya Shah&amp;rdquo;, that we added the model to catch.&lt;/p&gt;
&lt;p&gt;Look at the second row again. An identical name scored 0.76, not 0.99. We asked for &amp;ldquo;the same real entity&amp;rdquo;, &amp;ldquo;not merely a similar name&amp;rdquo;, and Jev priced in that a shared name is not proof of a shared identity. That is what you want from a decision about who someone is.&lt;/p&gt;
&lt;p&gt;Seven probes are not a benchmark, and we do not present them as one. What transfers is the method: log the probabilities on your own traffic, look for the gap, and place the threshold by what each kind of mistake costs. Then pin the model version if your route lets you: TypeSafe&amp;rsquo;s docs say to pin a version&amp;rsquo;s ID once you have tuned confidence thresholds against it (&lt;a href="https://docs.typesafe.ai/models" data-external&gt;TypeSafe docs&lt;/a&gt;). Cloudflare&amp;rsquo;s &lt;code&gt;typesafe/jev&lt;/code&gt; route, the one we call, has no version field, so it answers with whatever TypeSafe serves (&lt;a href="https://developers.cloudflare.com/ai/models/typesafe/jev/" data-external&gt;Cloudflare&lt;/a&gt;). Every response names the model that answered, so record it next to each decision and treat a change as a reason to measure again.&lt;/p&gt;
&lt;h3 id="use-the-probability-not-the-top-answer"&gt;Use the probability, not the top answer&lt;a class="heading-anchor" href="#use-the-probability-not-the-top-answer" aria-label="Link to this section"&gt;&lt;/a&gt;
&lt;/h3&gt;
&lt;p&gt;The most important line in our integration is a helper that refuses to use Jev&amp;rsquo;s top answer unless its probability clears the threshold. Real matches do not always clear it. In a case recorded in our code outside that probe set, Jev picked the right candidate at 0.63 with &amp;ldquo;none&amp;rdquo; holding 0.37. Taking the top answer would have merged, and this time it would have been right. We still do not merge at 0.63. A rule that merges at 63% will also merge the wrong candidate whenever Jev is 63% sure of it, and a wrong merge costs far more than the duplicate this leaves behind.&lt;/p&gt;
&lt;p&gt;If you take the top label and discard the probability, you have thrown away the one thing you were paying for. This is the same arithmetic as in &lt;a href="https://construct.computer/blog/agent-verification-gap/" data-external&gt;Nobody merges an email&lt;/a&gt;: delegating pays when the cost of checking plus the expected cost of a mistake is lower than doing the work yourself, and the model only shows up in the probability of a mistake. A generative judge hands you an answer. A decision model hands you an answer and that probability, so you can spend human attention where it is low instead of rereading everything.&lt;/p&gt;
&lt;h3 id="is-jevs-confidence-calibrated"&gt;Is Jev&amp;rsquo;s confidence calibrated?&lt;a class="heading-anchor" href="#is-jevs-confidence-calibrated" aria-label="Link to this section"&gt;&lt;/a&gt;
&lt;/h3&gt;
&lt;p&gt;Partly. TypeSafe trains for calibration and is explicit that calibration &amp;ldquo;does not guarantee that an individual answer is correct&amp;rdquo; (&lt;a href="https://docs.typesafe.ai/concepts/system-one" data-external&gt;TypeSafe docs&lt;/a&gt;). Independent tests agree the probabilities rank answers well and disagree with each other about how literally to read them:&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;OpenRouter found Jev &amp;ldquo;overestimated its own accuracy mid-range. Still, it ranks well.&amp;rdquo; At a confidence of 0.99 or more, which covered 58% of items, it was right 96.3% of the time (&lt;a href="https://openrouter.ai/blog/insights/jev-vs-claude-opus-5-classification/" data-external&gt;OpenRouter&lt;/a&gt;).&lt;/li&gt;
&lt;li&gt;A Towards Data Science test found answers in the 0.7 to 0.9 band right only 53% of the time, and answers at exactly 1.00 right 97.1% of the time (&lt;a href="https://towardsdatascience.com/jev-vs-llms-when-ai-moves-from-generation-to-decision-making/" data-external&gt;Towards Data Science&lt;/a&gt;).&lt;/li&gt;
&lt;li&gt;Primeline measured calibration error by question type: 0.012 for Noul, 0.086 for Choice, and 0.254 for Score, roughly a 20 times spread (&lt;a href="https://primeline.cc/blog/typesafe-jev-pre-registered-test" data-external&gt;Primeline&lt;/a&gt;).&lt;/li&gt;
&lt;/ul&gt;
&lt;p&gt;Treat Jev&amp;rsquo;s probabilities as a strong ranking, not as literal odds, and be most careful with Score questions. Our threshold does exactly that: it sits in a gap we observed rather than at a number we assumed. With a few hundred labeled examples from your own data you can go further and recalibrate the scores (&lt;a href="https://www.alexmolas.com/2026/09/23/jev-cant-be-calibrated.html" data-external&gt;Alex Molas&lt;/a&gt;).&lt;/p&gt;
&lt;h2 id="what-jev-gets-wrong"&gt;What Jev gets wrong&lt;a class="heading-anchor" href="#what-jev-gets-wrong" aria-label="Link to this section"&gt;&lt;/a&gt;
&lt;/h2&gt;
&lt;p&gt;TypeSafe publishes its own list of weak spots for &lt;code&gt;jev-1.13&lt;/code&gt; (&lt;a href="https://docs.typesafe.ai/model-jaggedness/jev-1.13" data-external&gt;TypeSafe docs&lt;/a&gt;). It reads instructions literally, &amp;ldquo;is not a calculator&amp;rdquo;, &amp;ldquo;does not count reliably&amp;rdquo;, struggles with date and time comparison and with multi-step indirection, degrades on large states full of irrelevant detail, and does not treat text in the state as hostile. Independent tests confirm the ones that matter most for agents:&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;&lt;strong&gt;Injected text moves the answer.&lt;/strong&gt; Primeline appended an &amp;ldquo;IMPORTANT INSTRUCTION FOR THE AI CLASSIFIER&amp;rdquo; line to an innocent reminder and watched its spam score rise from 0.04 to 0.66. Across 40 test pairs, 22.5% were misclassified outright (&lt;a href="https://primeline.cc/blog/typesafe-jev-pre-registered-test" data-external&gt;Primeline&lt;/a&gt;). A separate paper found injected content shifts probabilities but rarely makes Jev pick the attacker&amp;rsquo;s target, with adaptive attacks raising success from 1.8% to 3.5% (&lt;a href="https://arxiv.org/abs/2609.28613" data-external&gt;arXiv:2609.28613&lt;/a&gt;). For an agent that reads email and web pages, assume the state is attacker-controlled.&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Long, real-world state hurts.&lt;/strong&gt; Synthetic padding up to 125,000 characters cost Primeline nothing, but on real notes accuracy fell from 97.0% on the shortest quarter to 81.1% on the longest.&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Probability has to land somewhere.&lt;/strong&gt; Given 30 messages that fit none of the listed categories and no &amp;ldquo;none&amp;rdquo; option, Jev chose a listed category every time, at 0.99 confidence or more, in a test reported by &lt;a href="https://towardsdatascience.com/jev-vs-llms-when-ai-moves-from-generation-to-decision-making/" data-external&gt;Towards Data Science&lt;/a&gt;.&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Separate questions are not kept consistent.&lt;/strong&gt; In TypeSafe&amp;rsquo;s own docs, a question and its negation returned probabilities that added up to 1.19.&lt;/li&gt;
&lt;/ul&gt;
&lt;p&gt;These are the rules we wrote into our integration, and they follow directly from that list:&lt;/p&gt;
&lt;ol&gt;
&lt;li&gt;Never put Jev on an auth boundary, and never make it the only check before an irreversible action.&lt;/li&gt;
&lt;li&gt;Never ask it to compare dates, count, or do arithmetic. Compute those in code and put the result in the state.&lt;/li&gt;
&lt;li&gt;Always offer an explicit &amp;ldquo;none&amp;rdquo;.&lt;/li&gt;
&lt;li&gt;Ask one question per decision, and never infer the probability of &amp;ldquo;not x&amp;rdquo; from the probability of &amp;ldquo;x&amp;rdquo;.&lt;/li&gt;
&lt;li&gt;Keep the state small and relevant: the candidates, not the whole conversation.&lt;/li&gt;
&lt;li&gt;Treat a missing or malformed answer as no verdict, never as &amp;ldquo;no&amp;rdquo;, and keep the old path as a fallback. That includes a label you never offered: count it as no answer, not as &amp;ldquo;none&amp;rdquo;.&lt;/li&gt;
&lt;li&gt;Batch every question about the same state into one call.&lt;/li&gt;
&lt;li&gt;Record which model version answered, with the probability and what you did with it, so the threshold can be checked again later.&lt;/li&gt;
&lt;li&gt;Put anything a third party wrote in fields marked as untrusted, and tell every question to treat those fields as evidence, never as instructions.&lt;/li&gt;
&lt;/ol&gt;
&lt;h2 id="where-jev-fits-in-an-ai-agent-system-1-and-system-2"&gt;Where Jev fits in an AI agent: System 1 and System 2&lt;a class="heading-anchor" href="#where-jev-fits-in-an-ai-agent-system-1-and-system-2" aria-label="Link to this section"&gt;&lt;/a&gt;
&lt;/h2&gt;
&lt;p&gt;An AI employee makes two kinds of decisions. The big ones, what to do and how, belong to a language model that can plan, write, and use tools. The small ones happen dozens of times per task, and most of them are the shape Jev was built for: a bounded question over a small state, where a fast answer with a probability is worth more than a slow paragraph.&lt;/p&gt;
&lt;p&gt;Anthony Maio&amp;rsquo;s division of labor is the cleanest statement of the pattern we have seen: &amp;ldquo;A generative model drafts, plans, or explains. Jev supplies bounded semantic judgments. Code handles state, arithmetic, policy, permissions, and side effects. Humans take the ambiguous or high-risk cases&amp;rdquo; (&lt;a href="https://anthonymaio.substack.com/p/jev-the-language-model-that-wont" data-external&gt;Anthony Maio&lt;/a&gt;).&lt;/p&gt;
&lt;p&gt;This is how we rate the decision points in an always-on agent&amp;rsquo;s loop:&lt;/p&gt;
&lt;table&gt;
	&lt;thead&gt;
			&lt;tr&gt;
					&lt;th&gt;Decision&lt;/th&gt;
					&lt;th&gt;Jev question&lt;/th&gt;
					&lt;th&gt;Fit&lt;/th&gt;
					&lt;th&gt;What to design around&lt;/th&gt;
			&lt;/tr&gt;
	&lt;/thead&gt;
	&lt;tbody&gt;
			&lt;tr&gt;
					&lt;td&gt;Is this person or project one the agent already knows?&lt;/td&gt;
					&lt;td&gt;Choice, with &amp;ldquo;none&amp;rdquo;&lt;/td&gt;
					&lt;td&gt;In production at Construct&lt;/td&gt;
					&lt;td&gt;A wrong merge is hard to undo, so merge only above a measured threshold&lt;/td&gt;
			&lt;/tr&gt;
			&lt;tr&gt;
					&lt;td&gt;Is the task done, or should the agent keep going?&lt;/td&gt;
					&lt;td&gt;Noul&lt;/td&gt;
					&lt;td&gt;Good&lt;/td&gt;
					&lt;td&gt;The state includes tool output, which is untrusted, so stop when unsure&lt;/td&gt;
			&lt;/tr&gt;
			&lt;tr&gt;
					&lt;td&gt;Does this inbound email need the agent at all?&lt;/td&gt;
					&lt;td&gt;Choice: reply, file, or ignore&lt;/td&gt;
					&lt;td&gt;Good for skipping and labeling&lt;/td&gt;
					&lt;td&gt;Email is attacker-controlled, so the answer must never send anything&lt;/td&gt;
			&lt;/tr&gt;
			&lt;tr&gt;
					&lt;td&gt;Is this thread message addressed to the agent?&lt;/td&gt;
					&lt;td&gt;Noul&lt;/td&gt;
					&lt;td&gt;Good&lt;/td&gt;
					&lt;td&gt;Low stakes: a miss costs a slower reply&lt;/td&gt;
			&lt;/tr&gt;
			&lt;tr&gt;
					&lt;td&gt;Is there anything new since the last scheduled run?&lt;/td&gt;
					&lt;td&gt;Noul over a change list built in code&lt;/td&gt;
					&lt;td&gt;Promising&lt;/td&gt;
					&lt;td&gt;Jev is weak at dates, so work out what changed before you ask&lt;/td&gt;
			&lt;/tr&gt;
			&lt;tr&gt;
					&lt;td&gt;Which tool or skill fits this request?&lt;/td&gt;
					&lt;td&gt;Choice, up to 255 labels&lt;/td&gt;
					&lt;td&gt;Good as a suggestion&lt;/td&gt;
					&lt;td&gt;Let the planning model keep the final say&lt;/td&gt;
			&lt;/tr&gt;
			&lt;tr&gt;
					&lt;td&gt;Can this action run without asking the user?&lt;/td&gt;
					&lt;td&gt;Noul&lt;/td&gt;
					&lt;td&gt;Only as one layer&lt;/td&gt;
					&lt;td&gt;Never the only gate on something irreversible&lt;/td&gt;
			&lt;/tr&gt;
	&lt;/tbody&gt;
&lt;/table&gt;
&lt;p&gt;&lt;em&gt;Decisions in an AI employee&amp;rsquo;s loop, and how well each fits a decision model&lt;/em&gt;&lt;/p&gt;
&lt;p&gt;Two patterns come out of that table.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;The saving is the turns you do not run.&lt;/strong&gt; Comparing Jev&amp;rsquo;s price with a cheap language model&amp;rsquo;s misses where the money goes. When an email arrives or a schedule fires, the expensive part is the full agent turn that follows. A half-second check that says &amp;ldquo;this is a newsletter&amp;rdquo; or &amp;ldquo;nothing changed since the last run&amp;rdquo; lets the agent skip that turn entirely. TypeSafe&amp;rsquo;s skill-suggestion cookbook shows a similar effect inside a turn: with a Jev suggestion, an agent built on Claude Haiku 4.5 loaded the wrong skill 7.3% of the time instead of 16.8% (&lt;a href="https://docs.typesafe.ai/cookbooks/skill_suggestion" data-external&gt;TypeSafe docs&lt;/a&gt;).&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Escalate up, never sideways.&lt;/strong&gt; The obvious move is to send low-confidence decisions to a bigger model, and it works only if the fallback is actually better at the hard cases. OpenRouter sent Jev&amp;rsquo;s answers below 0.90 confidence to Claude Opus 5 and got within 0.4 points of Opus alone, with 76% of traffic never touching Opus, an upper bound because the cutoff was tuned on the test set (&lt;a href="https://openrouter.ai/blog/insights/jev-vs-claude-opus-5-classification/" data-external&gt;OpenRouter&lt;/a&gt;). A Towards Data Science test sent them to a weaker local model instead, and every cascade did worse than Jev alone: one &amp;ldquo;fixed 84 mistakes, but introduced 211 new ones&amp;rdquo; (&lt;a href="https://towardsdatascience.com/jev-vs-llms-when-ai-moves-from-generation-to-decision-making/" data-external&gt;Towards Data Science&lt;/a&gt;). Knowing which cases are hard is not the same as having something that can solve them. For irreversible actions, the right place to escalate is usually a person.&lt;/p&gt;
&lt;p&gt;&lt;em&gt;Read next: &lt;a href="https://construct.computer/blog/agent-task-half-life/" data-external&gt;Your agent has a half-life&lt;/a&gt;&lt;/em&gt;&lt;/p&gt;
&lt;h2 id="what-does-jev-cost-compared-with-an-llm"&gt;What does Jev cost compared with an LLM?&lt;a class="heading-anchor" href="#what-does-jev-cost-compared-with-an-llm" aria-label="Link to this section"&gt;&lt;/a&gt;
&lt;/h2&gt;
&lt;p&gt;At list prices, for a decision with about 400 tokens of input, and assuming a 20-token JSON answer from the models that generate one:&lt;/p&gt;
&lt;table&gt;
	&lt;thead&gt;
			&lt;tr&gt;
					&lt;th&gt;Model&lt;/th&gt;
					&lt;th&gt;Price per million tokens, input / output&lt;/th&gt;
					&lt;th&gt;Cost per 1,000 decisions&lt;/th&gt;
			&lt;/tr&gt;
	&lt;/thead&gt;
	&lt;tbody&gt;
			&lt;tr&gt;
					&lt;td&gt;Jev 1.13&lt;/td&gt;
					&lt;td&gt;$0.042 / free&lt;/td&gt;
					&lt;td&gt;$0.017&lt;/td&gt;
			&lt;/tr&gt;
			&lt;tr&gt;
					&lt;td&gt;Llama 4 Scout on Workers AI (our old judge)&lt;/td&gt;
					&lt;td&gt;$0.27 / $0.85&lt;/td&gt;
					&lt;td&gt;$0.125&lt;/td&gt;
			&lt;/tr&gt;
			&lt;tr&gt;
					&lt;td&gt;Claude Haiku 4.5&lt;/td&gt;
					&lt;td&gt;$1 / $5&lt;/td&gt;
					&lt;td&gt;$0.50&lt;/td&gt;
			&lt;/tr&gt;
			&lt;tr&gt;
					&lt;td&gt;Claude Sonnet 5.5&lt;/td&gt;
					&lt;td&gt;$2 / $10&lt;/td&gt;
					&lt;td&gt;$1.00&lt;/td&gt;
			&lt;/tr&gt;
	&lt;/tbody&gt;
&lt;/table&gt;
&lt;p&gt;&lt;em&gt;List-price cost per 1,000 decisions of about 400 input tokens, September 2026&lt;/em&gt;&lt;/p&gt;
&lt;p&gt;Two notes make the comparison fairer. Prompt caching cuts repeated input on Claude models to about a tenth of the price, but Haiku 4.5 only caches prompts of 4,096 tokens or more, so a short decision prompt gets no discount. Jev has no prompt caching at all, and Primeline saw identical requests billed in full every time, which barely matters at $0.042 per million tokens. Our old judge also needed up to three calls per entity, so its real cost was up to three times its row.&lt;/p&gt;
&lt;p&gt;The honest conclusion is less dramatic than the launch claim of &amp;ldquo;444.6x cheaper&amp;rdquo;. Independent tests put Jev at roughly 8 to 20 times cheaper per call than Haiku 4.5 and 10 to 20 times faster than strong language models on like-for-like classification, a few accuracy points behind the best of them (&lt;a href="https://primeline.cc/blog/typesafe-jev-pre-registered-test" data-external&gt;Primeline&lt;/a&gt;; &lt;a href="https://openrouter.ai/blog/insights/jev-vs-claude-opus-5-classification/" data-external&gt;OpenRouter&lt;/a&gt;). For one founder&amp;rsquo;s agent making a few hundred small decisions a day, the difference is cents a month either way. Price starts to matter at fleet scale or in tight loops, and the bigger saving is still the agent turns a cheap decision lets you skip.&lt;/p&gt;
&lt;p&gt;For us, cost was not the reason to switch. One call replaced up to three sequential ones, &amp;ldquo;none&amp;rdquo; became an explicit, scored answer, and we got a probability to put a threshold on.&lt;/p&gt;
&lt;h2 id="how-we-are-adding-more-jev-decisions"&gt;How we are adding more Jev decisions&lt;a class="heading-anchor" href="#how-we-are-adding-more-jev-decisions" aria-label="Link to this section"&gt;&lt;/a&gt;
&lt;/h2&gt;
&lt;p&gt;Entity matching was the first decision. Mapping every other place our agent makes a small judgment turned up about thirty more candidates, from &amp;ldquo;does this email need the agent?&amp;rdquo; to &amp;ldquo;did this scheduled check find anything new?&amp;rdquo;. Before moving any of them, we built one small runtime that every decision goes through. The lesson of the first week was that a decision you cannot see is a decision you cannot trust.&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;&lt;strong&gt;Every decision ships switched off.&lt;/strong&gt; A flag moves each one from off to shadow to live on its own, with its own threshold.&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Shadow before live.&lt;/strong&gt; In shadow the old path still decides, and Jev&amp;rsquo;s answer is recorded beside it. The old path&amp;rsquo;s answer becomes a free label, so the threshold comes from real traffic rather than a handful of probes.&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Records carry numbers, not text.&lt;/strong&gt; We keep the decision name, the label, the probability, what the product did, and which model answered, never the text Jev saw. Samples are deleted after 90 days.&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;The model version is recorded, not assumed.&lt;/strong&gt; Our route cannot pin a version, so each record names the version that answered, and a change raises an alarm. New decisions drop back to shadow until they are measured again.&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;A failing provider degrades to the old behaviour.&lt;/strong&gt; A circuit breaker stops calling Jev after repeated failures, and a rate gate keeps us well under TypeSafe&amp;rsquo;s published limit, 80 requests a second when we last checked (&lt;a href="https://docs.typesafe.ai/models" data-external&gt;TypeSafe docs&lt;/a&gt;). Every decision has a fallback, so an outage looks like the product before Jev.&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Third-party text is labelled as such.&lt;/strong&gt; Email bodies, web pages and other people&amp;rsquo;s messages go into fields named as untrusted, and every question says those fields are evidence, not instructions. In one small practitioner test, a single sentence like that cut targeted injection success to 1.3% (&lt;a href="https://github.com/Iskandeur/system1-system2" data-external&gt;Iskandeur&lt;/a&gt;). It is a mitigation, not a boundary, so a decision that reads untrusted text may only lower a priority or skip work, never grant anything.&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Decisions that read your conversations or email wait for the paperwork.&lt;/strong&gt; Sending your text to another provider is a privacy decision before it is an engineering one. TypeSafe is on our &lt;a href="https://construct.computer/sub-processors/" data-external&gt;sub-processors list&lt;/a&gt;, and the decisions that would send conversation or email text stay off until the contract review is done.&lt;/li&gt;
&lt;/ul&gt;
&lt;p&gt;The mapping also found plain bugs that had nothing to do with Jev. One was a keyword table that guessed which app a request needed: any request containing the letter x, such as &amp;ldquo;export as csv&amp;rdquo;, read as a request for Twitter, and &amp;ldquo;linear regression&amp;rdquo; read as a request for Linear. A keyword list is how you make a System One judgment without a System One model, and that is what it looks like when it fails.&lt;/p&gt;
&lt;h2 id="how-to-try-jev"&gt;How to try Jev&lt;a class="heading-anchor" href="#how-to-try-jev" aria-label="Link to this section"&gt;&lt;/a&gt;
&lt;/h2&gt;
&lt;ul&gt;
&lt;li&gt;&lt;strong&gt;Direct:&lt;/strong&gt; sign up at &lt;a href="https://typesafe.ai" data-external&gt;typesafe.ai&lt;/a&gt;. TypeSafe paused new signups on September 22 because of demand, and its CEO said on September 27 that signups had reopened, without free credits for new accounts (&lt;a href="https://x.com/CompleteSkeptic/status/2104338649999626397" data-external&gt;Diogo Almeida on X&lt;/a&gt;).&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Through a gateway, without a TypeSafe account:&lt;/strong&gt; OpenRouter (&lt;code&gt;typesafe/jev-1.13&lt;/code&gt;), Vercel AI Gateway (&lt;code&gt;typesafe-ai/jev&lt;/code&gt;), or Cloudflare Workers AI (&lt;code&gt;typesafe/jev&lt;/code&gt;). We call it through Cloudflare Workers AI and AI Gateway.&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;On Cloudflare, check the key alias:&lt;/strong&gt; the Workers AI binding only uses a provider key stored under the &lt;code&gt;default&lt;/code&gt; alias. With any other alias, calls quietly fall through to Cloudflare&amp;rsquo;s Unified Billing, which is capped at 200 requests a minute per gateway (&lt;a href="https://developers.cloudflare.com/ai-gateway/usage/worker-binding-methods/" data-external&gt;Cloudflare&lt;/a&gt;).&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;SDKs:&lt;/strong&gt; &lt;code&gt;typesafe-sdk&lt;/code&gt; for Python and &lt;code&gt;@typesafe-ai/sdk&lt;/code&gt; for JavaScript and TypeScript (&lt;a href="https://docs.typesafe.ai/introduction/quickstart" data-external&gt;TypeSafe quick start&lt;/a&gt;).&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Price:&lt;/strong&gt; $0.042 per million input tokens. Output tokens are counted but not billed. Several unofficial lookalike sites resell Jev access at several times that price, so check the price on typesafe.ai or your gateway before you pay.&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Version:&lt;/strong&gt; &lt;code&gt;jev-latest&lt;/code&gt; points at &lt;code&gt;jev-1.13.0&lt;/code&gt; today. If you tune thresholds, pin the version where your route allows it, and record the served version where it does not.&lt;/li&gt;
&lt;/ul&gt;
&lt;p&gt;Start with one decision that is frequent, runs in the background, is not an auth boundary, and already has a fallback. Log the probabilities for a week before you let them act on anything.&lt;/p&gt;
&lt;h2 id="what-most-jev-coverage-gets-wrong"&gt;What most Jev coverage gets wrong&lt;a class="heading-anchor" href="#what-most-jev-coverage-gets-wrong" aria-label="Link to this section"&gt;&lt;/a&gt;
&lt;/h2&gt;
&lt;ul&gt;
&lt;li&gt;&lt;strong&gt;&amp;ldquo;Jev never hallucinates.&amp;rdquo;&lt;/strong&gt; It cannot return a value outside your schema. It can return the wrong valid one, and TypeSafe&amp;rsquo;s own FAQ says it &amp;ldquo;can choose the wrong one&amp;rdquo; (&lt;a href="https://typesafe.ai" data-external&gt;TypeSafe&lt;/a&gt;).&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;&amp;ldquo;Jev is 67.8% accurate.&amp;rdquo;&lt;/strong&gt; That figure is agreement with the averaged answers of GPT-6 Astra and Claude Fable 5.1 on TypeSafe&amp;rsquo;s own four workflow evals, not accuracy against verified labels (&lt;a href="https://evals.typesafe.ai/" data-external&gt;TypeSafe evals&lt;/a&gt;).&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;&amp;ldquo;193.6x faster and 444.6x cheaper.&amp;rdquo;&lt;/strong&gt; On TypeSafe&amp;rsquo;s eval chart, those ratios line up with the slowest model, Claude Sonnet 5 at 78.1 seconds per case, and the most expensive, Claude Opus 5 at $0.1761 per case. Against GPT-5.6 Terra, which matched Jev&amp;rsquo;s agreement score, the same chart works out to about 25 times faster and 76 times cheaper.&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;&amp;ldquo;Jev has a 32k context window.&amp;rdquo;&lt;/strong&gt; It takes 64k tokens per request, of which 32k covers the state plus the longest question (&lt;a href="https://docs.typesafe.ai/models" data-external&gt;TypeSafe docs&lt;/a&gt;). Cloudflare&amp;rsquo;s model page lists 32,000 tokens for its route, so check the limit on the route you actually call.&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;&amp;ldquo;It is deterministic.&amp;rdquo;&lt;/strong&gt; TypeSafe designs for consistency rather than determinism. Our repeated calls moved by 0.02 at most, which is consistent, not identical.&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;&amp;ldquo;Just send the uncertain ones to a bigger model.&amp;rdquo;&lt;/strong&gt; Only if that model is better at the hard cases, as the escalation results above show.&lt;/li&gt;
&lt;/ul&gt;
&lt;h2 id="where-construct-fits"&gt;Where Construct fits&lt;a class="heading-anchor" href="#where-construct-fits" aria-label="Link to this section"&gt;&lt;/a&gt;
&lt;/h2&gt;
&lt;p&gt;Construct is an AI employee with its own computer: a workspace filesystem, long-term memory, a schedule, email, and workflows you can read before you run them. Jev runs in one place today, deciding whether a person or project your agent just learned about is one it already knows. Planning and writing run on general-purpose language models.&lt;/p&gt;
&lt;p&gt;Memory in Construct is inspectable and correctable, so if a merge pulls the wrong facts together, you can see which ones and correct or forget them. &lt;a href="https://construct.computer/blog/ai-agent-memory/" data-external&gt;AI agent memory&lt;/a&gt; covers how that works. Two honest boundaries: the &amp;ldquo;is this task done?&amp;rdquo; check and email triage are written as Jev decisions, but both stay switched off until shadow data shows they beat what they would replace, and Construct does not currently insert a mandatory approval gate before every external side effect, so steps that send mail or move money still need supervision before they run unattended.&lt;/p&gt;
&lt;p&gt;If the reliability argument is new to you, start with &lt;a href="https://construct.computer/blog/agent-task-half-life/" data-external&gt;your agent has a half-life&lt;/a&gt; and &lt;a href="https://construct.computer/blog/agent-verification-gap/" data-external&gt;nobody merges an email&lt;/a&gt;. For the infrastructure underneath, read &lt;a href="https://construct.computer/blog/running-ai-agents-on-cloudflare-not-vms/" data-external&gt;how our agents get computers we mostly do not pay for&lt;/a&gt;, and if the category itself is new, begin with &lt;a href="https://construct.computer/blog/what-is-an-ai-employee/" data-external&gt;what is an AI employee&lt;/a&gt;.&lt;/p&gt;
&lt;hr&gt;
&lt;p&gt;&lt;em&gt;Originally posted on &lt;a href="https://construct.computer/blog/jev-ai-agents/" data-external&gt;construct.computer&lt;/a&gt;.&lt;/em&gt;&lt;/p&gt;</content:encoded></item><item><title>Nobody merges an email</title><link>https://ankush.one/blogs/agent-verification-gap/</link><pubDate>Tue, 22 Sep 2026 00:00:00 +0000</pubDate><author>ankush4singh@gmail.com (Ankush Singh)</author><guid>https://ankush.one/blogs/agent-verification-gap/</guid><category>ai-agents</category><category>construct</category><description>Agents made work cheap to produce and no cheaper to check. Why software absorbed the flood, why the rest of the business did not, and the five things that make agent work checkable.</description><content:encoded>&lt;p&gt;It is 5:40 on a Friday and the agent has just written nine client updates in four minutes, sitting there quietly pleased with itself. Eight of them are fine. One of them has told a client something that is not true, and because nothing in the pipeline can tell you which one, you read all nine. Forty minutes later you send eight, rewrite one, and work out that the time you saved on the writing came straight back out of your evening. The agent did not save you thirty six minutes. It moved them, out of writing, which you are decent at, and into checking, which nobody is good at and nobody enjoys.&lt;/p&gt;
&lt;p&gt;Every tool you actually work in has an undo. Git has revert, your editor has a keystroke, staging has the comfort of breaking somewhere that does not count, and Gmail gives you thirty seconds. Nobody merges an email: there is no diff to skim, no test suite to run it past, no staging inbox to try it in, and no way back once it is sitting in somebody&amp;rsquo;s inbox, so the entire cost of being wrong lands on one person reading carefully. That gap, and not the model, is why your agent is trusted with code and kept well away from the account manager&amp;rsquo;s inbox. It is also why most agent pilots quietly stall.&lt;/p&gt;
&lt;h2 id="the-bottleneck-moved-and-the-tooling-stayed-put"&gt;The bottleneck moved and the tooling stayed put&lt;a class="heading-anchor" href="#the-bottleneck-moved-and-the-tooling-stayed-put" aria-label="Link to this section"&gt;&lt;/a&gt;
&lt;/h2&gt;
&lt;p&gt;Software is the one place where a lot of people pointed agents at real work and then measured what happened. The measurements are not subtle.&lt;/p&gt;
&lt;p&gt;Faros AI compared two years of telemetry from 22,000 developers across 4,000 teams, looking at each organization at its lowest and highest AI adoption. Output went up. So did everything that output costs you.&lt;/p&gt;
&lt;figure class="doc-figure"&gt;&lt;img src="https://ankush.one/media/construct/agent-verification-gap/acceleration-whiplash.svg" alt="Horizontal bar chart of percentage change between each organization’s lowest and highest AI adoption. Tasks completed per developer up 33.7 percent and epics completed up 66 percent, against bugs per developer up 54 percent, median time to first review up 156.6 percent, incidents per pull request up 242.7 percent, median review time up 441.5 percent, and code churn up 861 percent." title="The two shortest bars are the benefit. Everything longer is what the benefit cost. From Faros AI&amp;#39;s 2026 report, The Acceleration Whiplash." loading="lazy" decoding="async"&gt;&lt;figcaption&gt;The two shortest bars are the benefit. Everything longer is what the benefit cost. From Faros AI&amp;#39;s 2026 report, The Acceleration Whiplash.&lt;/figcaption&gt;&lt;/figure&gt;&lt;p&gt;A third more tasks per developer, and five times the median review time (&lt;a href="https://www.faros.ai/research/ai-acceleration-whiplash" data-external&gt;Faros AI, 2026&lt;/a&gt;). That is not a tooling problem or a model problem. That is a system that got very good at producing work and no better at absorbing it.&lt;/p&gt;
&lt;p&gt;LinearB looked at the same squeeze from the other end, across 8.1 million pull requests from 4,800 teams in 42 countries, and found the detail that gives the game away.&lt;/p&gt;
&lt;figure class="doc-figure"&gt;&lt;img src="https://ankush.one/media/construct/agent-verification-gap/pickup-vs-review.svg" alt="Grouped bar chart in hours per pull request. Waiting for a reviewer to start: 3.3 hours without AI against more than 16 hours for AI-assisted work. Actually being reviewed: 4.2 hours without AI against 3.2 hours for AI-assisted work." title="AI-assisted work waits more than five times longer for someone to start reviewing it, then takes slightly less time to review than human work once they do. From LinearB&amp;#39;s 2026 Engineering Benchmarks Report." loading="lazy" decoding="async"&gt;&lt;figcaption&gt;AI-assisted work waits more than five times longer for someone to start reviewing it, then takes slightly less time to review than human work once they do. From LinearB&amp;#39;s 2026 Engineering Benchmarks Report.&lt;/figcaption&gt;&lt;/figure&gt;&lt;p&gt;Read those two bars again. Once a reviewer actually starts, the AI-assisted change is &lt;em&gt;faster&lt;/em&gt; to get through than the human one: 194 minutes against 252 (&lt;a href="https://linearb.io/blog/8-million-prs-engineering-productivity" data-external&gt;LinearB, 2026&lt;/a&gt;). The delay is not the reading. The delay is everything before the reading: the queue, the size of the thing, the sinking feeling, the decision to start.&lt;/p&gt;
&lt;p&gt;The expensive part of verification was never the mechanical act of checking. It is the willingness to be the person who signs off.&lt;/p&gt;
&lt;h2 id="why-software-absorbed-the-flood-anyway"&gt;Why software absorbed the flood anyway&lt;a class="heading-anchor" href="#why-software-absorbed-the-flood-anyway" aria-label="Link to this section"&gt;&lt;/a&gt;
&lt;/h2&gt;
&lt;p&gt;Engineering took that hit and mostly kept moving, and it is worth being precise about why, because the reason is not talent or process maturity. It is that code is the most verifiable artifact a company produces, and it has been getting more verifiable for forty years.&lt;/p&gt;
&lt;figure class="doc-figure"&gt;&lt;img src="https://ankush.one/media/construct/agent-verification-gap/verification-stack.svg" alt="Two columns comparing five verification questions. Shipping code answers all five: git diff shows what changed, staging tries it without consequences, CI and tests have a machine check it, a pull request records who approved, and git revert undoes it. Everything else agents do answers none of them: you read it all again, there is no staging inbox, nothing runs, approval is a chat message at best, and sending is final." title="Every finished piece of work raises the same five questions. Software spent four decades building cheap answers to all of them. Nobody built the other column." loading="lazy" decoding="async"&gt;&lt;figcaption&gt;Every finished piece of work raises the same five questions. Software spent four decades building cheap answers to all of them. Nobody built the other column.&lt;/figcaption&gt;&lt;/figure&gt;&lt;p&gt;A diff is a strange and wonderful thing. It turns &amp;ldquo;read this 400 line change&amp;rdquo; into &amp;ldquo;read these 40 lines that differ&amp;rdquo;, which is a reduction of an order of magnitude in what a human has to hold in their head. Tests turn &amp;ldquo;is it correct&amp;rdquo; into a boolean somebody else already thought about. Staging turns &amp;ldquo;will this break&amp;rdquo; into &amp;ldquo;it did not break over there&amp;rdquo;. A pull request turns approval into a record with a name on it. Revert turns a mistake into an inconvenience.&lt;/p&gt;
&lt;p&gt;None of that exists for the work most of a company actually does.&lt;/p&gt;
&lt;p&gt;There is no diff for an email. There is no staging environment for a CRM. No test suite asserts that the invoice you generated has the right billing period on it, and no button unsends the reply to a customer.&lt;/p&gt;
&lt;p&gt;So when an agent produces that kind of work, the verification cost lands entirely on a human with no instruments, and it lands per unit of output. Double the output, double the checking. That is the whole story of why the demo felt like magic and the rollout felt like a second job.&lt;/p&gt;
&lt;p&gt;&lt;em&gt;Read next: &lt;a href="https://construct.computer/blog/agent-task-half-life/" data-external&gt;Your agent has a half-life&lt;/a&gt;&lt;/em&gt;&lt;/p&gt;
&lt;h2 id="the-check-has-to-be-cheaper-than-the-doing"&gt;The check has to be cheaper than the doing&lt;a class="heading-anchor" href="#the-check-has-to-be-cheaper-than-the-doing" aria-label="Link to this section"&gt;&lt;/a&gt;
&lt;/h2&gt;
&lt;p&gt;Here is the arithmetic nobody puts in the pitch deck.&lt;/p&gt;
&lt;p&gt;Let &lt;strong&gt;D&lt;/strong&gt; be what it costs you to do the job yourself. Let &lt;strong&gt;C&lt;/strong&gt; be what it costs to check the agent&amp;rsquo;s version. Let &lt;strong&gt;p&lt;/strong&gt; be the chance it is wrong in a way you would care about, and &lt;strong&gt;R&lt;/strong&gt; the cost when a wrong one gets through: the apology, the refund, the client who stops replying.&lt;/p&gt;
&lt;p&gt;Delegating pays when:&lt;/p&gt;
&lt;div class="highlight"&gt;&lt;pre tabindex="0" class="chroma"&gt;&lt;code class="language-text" data-lang="text"&gt;&lt;span class="line"&gt;&lt;span class="cl"&gt;C + p x R &amp;lt; D
&lt;/span&gt;&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;&lt;/div&gt;&lt;p&gt;The model only appears inside &lt;strong&gt;p&lt;/strong&gt;. Everything else in that inequality is a property of your tooling, and it is the part every agent roadmap ignores.&lt;/p&gt;
&lt;p&gt;Run it on the Friday nine. Writing one yourself is about six minutes, so D is 54 minutes. Reading nine drafts closely enough to put your name on them is about four minutes each, so C is 36, which is near enough the forty you actually spent. The saving is whatever is left, and it is thin. Note where it went, too: one draft was wrong, and you paid to check nine, because nothing in the pipeline could tell you which. Now make the drafts slightly harder to trust, so you reread the source thread for each one, and C passes D. The agent is now a net loss at 100% quality, because quality was never what you were paying for.&lt;/p&gt;
&lt;table&gt;
	&lt;thead&gt;
			&lt;tr&gt;
					&lt;th&gt;Work the agent finished&lt;/th&gt;
					&lt;th&gt;What you check it against&lt;/th&gt;
					&lt;th&gt;Cost to check&lt;/th&gt;
					&lt;th&gt;Cost if a bad one gets through&lt;/th&gt;
			&lt;/tr&gt;
	&lt;/thead&gt;
	&lt;tbody&gt;
			&lt;tr&gt;
					&lt;td&gt;A code change&lt;/td&gt;
					&lt;td&gt;A diff, a test run, a staging deploy&lt;/td&gt;
					&lt;td&gt;Minutes, tool-assisted&lt;/td&gt;
					&lt;td&gt;A revert&lt;/td&gt;
			&lt;/tr&gt;
			&lt;tr&gt;
					&lt;td&gt;Nine client emails&lt;/td&gt;
					&lt;td&gt;The nine source threads, from memory&lt;/td&gt;
					&lt;td&gt;Nearly the cost of writing them&lt;/td&gt;
					&lt;td&gt;A relationship&lt;/td&gt;
			&lt;/tr&gt;
			&lt;tr&gt;
					&lt;td&gt;200 CRM records updated&lt;/td&gt;
					&lt;td&gt;Nothing. You spot check twelve and hope&lt;/td&gt;
					&lt;td&gt;Unbounded, so people skip it&lt;/td&gt;
					&lt;td&gt;A quarter of bad data&lt;/td&gt;
			&lt;/tr&gt;
	&lt;/tbody&gt;
&lt;/table&gt;
&lt;p&gt;&lt;em&gt;What checking costs across three kinds of agent work&lt;/em&gt;&lt;/p&gt;
&lt;p&gt;Notice what happens in the bottom row. When C is unbounded, people do not pay it. They either check nothing and silently accept &lt;strong&gt;p x R&lt;/strong&gt;, or they stop using the agent for that work. Both look like &amp;ldquo;the agent did not work out&amp;rdquo;. Neither is about the agent.&lt;/p&gt;
&lt;p&gt;This is the same shape as &lt;a href="https://construct.computer/blog/agent-task-half-life/" data-external&gt;the half-life problem&lt;/a&gt;: a per-step property that looks fine up close and decides everything at scale. There, it was reliability compounding over a long run. Here, it is verification cost compounding over volume. Both get fixed by changing the structure of the run, not the intelligence inside it.&lt;/p&gt;
&lt;h2 id="nobody-is-measuring-this-including-you"&gt;Nobody is measuring this, including you&lt;a class="heading-anchor" href="#nobody-is-measuring-this-including-you" aria-label="Link to this section"&gt;&lt;/a&gt;
&lt;/h2&gt;
&lt;p&gt;The obvious objection is that if verification were really eating the gains, people would notice.&lt;/p&gt;
&lt;p&gt;They do not notice. That is the best-documented part of this whole argument.&lt;/p&gt;
&lt;p&gt;METR ran a randomized controlled trial with 16 experienced open-source developers on 246 real tasks in repositories they already maintained. Before starting, the developers forecast that AI tools would make them 24% faster. Afterwards, having done the work, they believed they had been 20% faster. Measured, they took 19% longer.&lt;/p&gt;
&lt;figure class="doc-figure"&gt;&lt;img src="https://ankush.one/media/construct/agent-verification-gap/metr-perception-gap.svg" alt="Diverging bar chart. Before the study, developers forecast being 24 percent faster with AI. After the study they believed they had been 20 percent faster. The measured result was 19 percent slower." title="A 39 point gap between what it felt like and what happened, from people with hundreds of hours of prompting experience. From METR&amp;#39;s July 2025 randomized controlled trial." loading="lazy" decoding="async"&gt;&lt;figcaption&gt;A 39 point gap between what it felt like and what happened, from people with hundreds of hours of prompting experience. From METR&amp;#39;s July 2025 randomized controlled trial.&lt;/figcaption&gt;&lt;/figure&gt;&lt;p&gt;These were not novices fumbling with a new toy. Nearly all of them had dozens to hundreds of hours of prior experience prompting models (&lt;a href="https://metr.org/blog/2025-07-10-early-2025-ai-experienced-os-dev-study/" data-external&gt;METR, July 2025&lt;/a&gt;). The generation felt fast, because it was. The review, the correction, the second pass to understand what had been written for them: that time was real and it did not register.&lt;/p&gt;
&lt;p&gt;Now scale that error up to a company. MIT&amp;rsquo;s Project NANDA, working from more than 300 publicly disclosed deployments, 52 interviews, and 153 survey responses, put the share of enterprise GenAI pilots with no measurable impact on the P&amp;amp;L at roughly 95% (&lt;a href="https://mlq.ai/media/quarterly_decks/v0.1_State_of_AI_in_Business_2025_Report.pdf" data-external&gt;MIT NANDA, 2025&lt;/a&gt;). People have argued about that methodology ever since. Nobody has argued that the number is small. Gartner expects more than 40% of agentic projects to be canceled by the end of 2027, citing costs and unclear value (&lt;a href="https://www.gartner.com/en/newsroom/press-releases/2025-06-25-gartner-predicts-over-40-percent-of-agentic-ai-projects-will-be-canceled-by-end-of-2027" data-external&gt;Gartner, June 2025&lt;/a&gt;). DORA, surveying nearly 5,000 professionals, landed on the phrase that ties it together: AI is an amplifier (&lt;a href="https://dora.dev/dora-report-2025/" data-external&gt;DORA, 2025&lt;/a&gt;). It makes a system with good feedback loops better and a system without them worse, faster.&lt;/p&gt;
&lt;p&gt;Most business processes have no feedback loops at all. They have a person who would notice eventually.&lt;/p&gt;
&lt;p&gt;Agents rarely fail an evaluation. They fail an audit nobody ran.&lt;/p&gt;
&lt;h2 id="five-things-that-make-agent-work-checkable"&gt;Five things that make agent work checkable&lt;a class="heading-anchor" href="#five-things-that-make-agent-work-checkable" aria-label="Link to this section"&gt;&lt;/a&gt;
&lt;/h2&gt;
&lt;p&gt;If verification cost is the constraint, then the interesting engineering is not in the agent loop. It is in everything the agent leaves behind. Five properties do most of the work, and none of them require a better model.&lt;/p&gt;
&lt;h3 id="artifacts-not-transcripts"&gt;Artifacts, not transcripts&lt;a class="heading-anchor" href="#artifacts-not-transcripts" aria-label="Link to this section"&gt;&lt;/a&gt;
&lt;/h3&gt;
&lt;p&gt;A chat transcript is the worst possible verification surface: linear, long, and it buries the output inside the reasoning. A file is the best one. You can open it, skim it, diff it against last month&amp;rsquo;s, hand it to someone else, and know at a glance whether the thing exists yet.&lt;/p&gt;
&lt;p&gt;The rule is that finished work should be an object you can point at, not a passage you have to read to the end of.&lt;/p&gt;
&lt;h3 id="a-draft-state-before-anything-irreversible"&gt;A draft state before anything irreversible&lt;a class="heading-anchor" href="#a-draft-state-before-anything-irreversible" aria-label="Link to this section"&gt;&lt;/a&gt;
&lt;/h3&gt;
&lt;p&gt;This is the missing diff, and it is mostly a product decision rather than a research problem. Generate the email, do not send it. Compute the 200 CRM updates, write them as a table first. Produce the invoice as a file before it goes to the customer.&lt;/p&gt;
&lt;p&gt;A draft turns an irreversible action into a reviewable artifact, which moves it out of the unbounded row of that table and into the cheap one. The work is identical. The verification cost is not.&lt;/p&gt;
&lt;h3 id="provenance-on-anything-the-agent-believes"&gt;Provenance on anything the agent believes&lt;a class="heading-anchor" href="#provenance-on-anything-the-agent-believes" aria-label="Link to this section"&gt;&lt;/a&gt;
&lt;/h3&gt;
&lt;p&gt;Half of checking is not &amp;ldquo;is this right&amp;rdquo;, it is &amp;ldquo;where did this come from&amp;rdquo;. An agent that says the client&amp;rsquo;s renewal is in March is unverifiable. An agent that says the renewal is in March, learned from the contract PDF you uploaded on 14 August, can be checked in four seconds.&lt;/p&gt;
&lt;p&gt;That is why &lt;a href="https://construct.computer/blog/ai-agent-memory/" data-external&gt;agent memory needs provenance and correction&lt;/a&gt; rather than a vector blob. Memory you cannot audit is a claim you have to re-derive, which is verification cost with extra steps.&lt;/p&gt;
&lt;h3 id="a-bounded-record-of-what-it-touched"&gt;A bounded record of what it touched&lt;a class="heading-anchor" href="#a-bounded-record-of-what-it-touched" aria-label="Link to this section"&gt;&lt;/a&gt;
&lt;/h3&gt;
&lt;p&gt;Not the full trace. Nobody reads the full trace. What you need is the ledger: what did it touch, when, and why. It is the difference between a stack trace and a receipt.&lt;/p&gt;
&lt;p&gt;The test is whether someone can answer &amp;ldquo;what did it actually do yesterday&amp;rdquo; in under a minute without reading a single model output.&lt;/p&gt;
&lt;h3 id="reversibility-where-it-exists-a-human-where-it-does-not"&gt;Reversibility where it exists, a human where it does not&lt;a class="heading-anchor" href="#reversibility-where-it-exists-a-human-where-it-does-not" aria-label="Link to this section"&gt;&lt;/a&gt;
&lt;/h3&gt;
&lt;p&gt;Some actions can be undone, and those should be cheap to approve. Some cannot, and pretending otherwise is how teams get burned. Sending mail, moving money, posting publicly, deleting records: for those the honest answer is that a person confirms, and the product&amp;rsquo;s job is to make that confirmation a two second glance rather than an investigation.&lt;/p&gt;
&lt;h2 id="how-construct-is-built-around-the-check"&gt;How Construct is built around the check&lt;a class="heading-anchor" href="#how-construct-is-built-around-the-check" aria-label="Link to this section"&gt;&lt;/a&gt;
&lt;/h2&gt;
&lt;p&gt;We build &lt;a href="https://construct.computer/blog/ai-employee/" data-external&gt;an AI employee&lt;/a&gt; with a real computer, so this is the problem we live in. A few things follow directly from the argument above.&lt;/p&gt;
&lt;p&gt;Work lands in a workspace filesystem as files, not as messages in a thread. That is the artifacts rule, and it is also why a run that dies partway through leaves usable output behind rather than a transcript to excavate.&lt;/p&gt;
&lt;p&gt;Every action goes into Activity: what it affected, when it ran, and why. That is the receipt, not the stack trace, and it exists so the answer to &amp;ldquo;what did it do overnight&amp;rdquo; is a scroll rather than a project. Alongside it you get the files it produced, the tool records, and a read-only terminal transcript when you want to go deeper.&lt;/p&gt;
&lt;p&gt;Memory is inspectable and correctable, with provenance and temporal context, so a claim can be traced to where it came from and fixed in place instead of argued with. Repeatable work becomes a saved workflow you can read before you schedule it. Current workflows are linear, which is a real limitation and also the reason a workflow is something you can hold in your head.&lt;/p&gt;
&lt;p&gt;And the part we have not finished: Construct does not currently insert a mandatory approval gate before every external side effect. Drafts, interruption mid-run, and Activity give you a lot, but if a step sends a customer email or moves money, that step still needs supervision before you let it run unattended. We would rather say that plainly than let someone discover it on a Sunday.&lt;/p&gt;
&lt;h2 id="the-next-10x-is-not-in-the-model"&gt;The next 10x is not in the model&lt;a class="heading-anchor" href="#the-next-10x-is-not-in-the-model" aria-label="Link to this section"&gt;&lt;/a&gt;
&lt;/h2&gt;
&lt;p&gt;Every agent product is currently competing on the part that got commoditized. Producing plausible work is close to free and getting freer.&lt;/p&gt;
&lt;p&gt;The scarce thing is a human&amp;rsquo;s willingness to sign their name under it, and that is bought with diffs, drafts, receipts, provenance, and undo. Software has had those for forty years and, funnily enough, software is the one place agents are actually working.&lt;/p&gt;
&lt;p&gt;The rest of the company is waiting for someone to build them the other column.&lt;/p&gt;
&lt;hr&gt;
&lt;p&gt;&lt;em&gt;Originally posted on &lt;a href="https://construct.computer/blog/agent-verification-gap/" data-external&gt;construct.computer&lt;/a&gt;.&lt;/em&gt;&lt;/p&gt;</content:encoded></item><item><title>Your agent has a half-life</title><link>https://ankush.one/blogs/agent-task-half-life/</link><pubDate>Tue, 11 Aug 2026 00:00:00 +0000</pubDate><author>ankush4singh@gmail.com (Ankush Singh)</author><guid>https://ankush.one/blogs/agent-task-half-life/</guid><category>ai-agents</category><category>construct</category><description>Why AI agents keep failing on long multi-step jobs: a 95% reliable agent finishes 48 steps 8.5% of the time. The fix is a resumable run, not a better model.</description><content:encoded>&lt;p&gt;Valve named a game &lt;em&gt;Half-Life&lt;/em&gt; after the physics joke hiding in plain sight: how long something lasts before half of it is gone. Your agent has one of those. It just does not get a hazmat suit or a crowbar out of the deal.&lt;/p&gt;
&lt;p&gt;Take an agent that gets 95% of its individual steps right. That is a good agent. Now give it a ten-step job.&lt;/p&gt;
&lt;p&gt;It finishes six times out of ten. The other four times, it fails halfway. Literally. A half-life.&lt;/p&gt;
&lt;p&gt;Nothing broke. No step regressed. 0.95 to the tenth power is 0.60, and that is the whole story. Reliability that looks excellent per step is mediocre per job, and it gets worse in a way most people do not feel until they schedule something and walk away.&lt;/p&gt;
&lt;p&gt;If you have been asking why your agent keeps failing halfway through a multi-step job, this is usually the answer, and it is not the one people go looking for.&lt;/p&gt;
&lt;h2 id="why-your-agent-keeps-failing-on-long-tasks"&gt;Why your agent keeps failing on long tasks&lt;a class="heading-anchor" href="#why-your-agent-keeps-failing-on-long-tasks" aria-label="Link to this section"&gt;&lt;/a&gt;
&lt;/h2&gt;
&lt;p&gt;The tempting explanation is that long tasks are harder tasks. Longer means more ambiguity, more chances to misread the goal, more compounding confusion.&lt;/p&gt;
&lt;p&gt;That explanation is wrong, and there is a clean piece of work showing why. In &amp;ldquo;Is there a half-life for the success rates of AI agents?&amp;rdquo;, Toby Ord shows that agent performance across task lengths is explained by an extremely simple model: a constant rate of failing during each minute a human would take to do the task (&lt;a href="https://arxiv.org/abs/2505.05115" data-external&gt;Ord, arXiv:2505.05115&lt;/a&gt;). Not accumulating confusion. A flat hazard rate, ticking the entire time the job runs.&lt;/p&gt;
&lt;p&gt;Physicists already have a letter for that rate. It is λ, the decay constant in N(t) = N₀·e^(−λt), and the half-life is just t½ = ln(2) / λ. Valve put the same λ on Gordon Freeman&amp;rsquo;s suit and in the &lt;em&gt;Half-Life&lt;/em&gt; logo for the same reason: it is the constant that decides how long something lasts before it is gone. Ord&amp;rsquo;s result is that your agent has one too.&lt;/p&gt;
&lt;p&gt;That gives every agent a half-life: the task duration at which its success probability hits 50%. Longer tasks fail, in Ord&amp;rsquo;s framing, because they contain increasingly large sets of subtasks where failing any one fails the whole thing. Raise λ (worse per-minute reliability) or raise exposure (longer jobs), and the surviving fraction falls the same way a sample of isotope does. You do not need a crowbar for this. You need less exposure.&lt;/p&gt;
&lt;p&gt;This reframes the problem usefully. If failure were about difficulty, you would fix it with a smarter model. If failure is a rate per unit of exposure, you fix it by reducing exposure. Those are very different engineering programs, and almost everyone is running the first one.&lt;/p&gt;
&lt;h2 id="why-multi-step-agent-workflows-fail-48-steps-85-success"&gt;Why multi-step agent workflows fail: 48 steps, 8.5% success&lt;a class="heading-anchor" href="#why-multi-step-agent-workflows-fail-48-steps-85-success" aria-label="Link to this section"&gt;&lt;/a&gt;
&lt;/h2&gt;
&lt;p&gt;Ten steps is a toy. Real recurring business work is worse, because it loops.&lt;/p&gt;
&lt;p&gt;Here is a job we hear about constantly from small agencies: monthly client reporting. Eight clients. For each one, pull analytics, pull ad spend, pull the CRM&amp;rsquo;s deal movement, write the summary, render it into the client&amp;rsquo;s template, email it. Six steps, eight times. Forty-eight steps.&lt;/p&gt;
&lt;table&gt;
	&lt;thead&gt;
			&lt;tr&gt;
					&lt;th&gt;Per-step reliability&lt;/th&gt;
					&lt;th&gt;One 48-step run finishes&lt;/th&gt;
			&lt;/tr&gt;
	&lt;/thead&gt;
	&lt;tbody&gt;
			&lt;tr&gt;
					&lt;td&gt;99%&lt;/td&gt;
					&lt;td&gt;62%&lt;/td&gt;
			&lt;/tr&gt;
			&lt;tr&gt;
					&lt;td&gt;95%&lt;/td&gt;
					&lt;td&gt;8.5%&lt;/td&gt;
			&lt;/tr&gt;
			&lt;tr&gt;
					&lt;td&gt;90%&lt;/td&gt;
					&lt;td&gt;0.6%&lt;/td&gt;
			&lt;/tr&gt;
	&lt;/tbody&gt;
&lt;/table&gt;
&lt;p&gt;A 95% agent, which is a genuinely capable agent, completes that job start to finish about one time in twelve. And the failure mode is the ugly one: it dies at client six, having already emailed five reports, and you cannot tell what state anything is in without reading the whole transcript. One continuous experiment, cascading into a mess you did not budget for. Black Mesa energy, spreadsheet edition.&lt;/p&gt;
&lt;p&gt;So people conclude agents do not work for real operations. What actually does not work is running forty-eight steps as one uninterrupted bet.&lt;/p&gt;
&lt;h2 id="will-a-better-model-fix-agent-reliability"&gt;Will a better model fix agent reliability?&lt;a class="heading-anchor" href="#will-a-better-model-fix-agent-reliability" aria-label="Link to this section"&gt;&lt;/a&gt;
&lt;/h2&gt;
&lt;p&gt;There is a genuine counterargument: models are getting better at long tasks fast. METR measures the 50% time horizon, the task duration at which an agent is predicted to succeed half the time, using how long human experts need for the same work. That horizon has been doubling roughly every seven months for six years, with recent data suggesting faster (&lt;a href="https://metr.org/blog/2025-03-19-measuring-ai-ability-to-complete-long-tasks/" data-external&gt;METR, measuring AI ability to complete long tasks&lt;/a&gt;; &lt;a href="https://metr.org/time-horizons/" data-external&gt;METR, time horizons&lt;/a&gt;).&lt;/p&gt;
&lt;p&gt;That is real progress and it is not slowing. It is also a terrible thing to plan a business process around. The horizon is defined at 50% reliability, which is a coin flip, and the doubling curve tells you nothing about next Tuesday&amp;rsquo;s client reports. Gartner predicts that over 40% of agentic AI projects will be canceled by the end of 2027, citing escalating costs, unclear business value, and inadequate risk controls (&lt;a href="https://www.gartner.com/en/newsroom/press-releases/2025-06-25-gartner-predicts-over-40-percent-of-agentic-ai-projects-will-be-canceled-by-end-of-2027" data-external&gt;Gartner, June 2025&lt;/a&gt;). A large share of those projects will die from exactly this: a demo that worked once, scheduled weekly, quietly failing most weeks. Unforeseen consequences, delivered on a cron.&lt;/p&gt;
&lt;p&gt;Waiting for the model is a strategy with no delivery date. Shortening the run is available now.&lt;/p&gt;
&lt;h2 id="how-to-fix-an-agent-that-fails-partway-cut-the-job-at-the-seams"&gt;How to fix an agent that fails partway: cut the job at the seams&lt;a class="heading-anchor" href="#how-to-fix-an-agent-that-fails-partway-cut-the-job-at-the-seams" aria-label="Link to this section"&gt;&lt;/a&gt;
&lt;/h2&gt;
&lt;p&gt;Split those forty-eight steps into eight independent six-step runs and the arithmetic inverts.&lt;/p&gt;
&lt;p&gt;At 95% per step, one client&amp;rsquo;s six-step run finishes 73.5% of the time. So a pass over eight clients lands roughly six of them, and the two that failed are named, isolated, and retryable. Run the retry and you are at 93% of the work done, with the remaining gap being two specific clients you can look at rather than one opaque dead job.&lt;/p&gt;
&lt;p&gt;Same model. Same per-step reliability. Same total work. The only change is where the run is allowed to stop and be resumed. You have not lowered λ. You have lowered how long any one run has to survive it.&lt;/p&gt;
&lt;p&gt;The catch is that this only works if state survives between the pieces. A six-step run that has to rebuild its own context from scratch, re-derive what the client&amp;rsquo;s template looks like, re-fetch what it already fetched, has not been shortened. It has been shortened and then padded back out. Checkpointing is only cheap when the checkpoint is written somewhere durable.&lt;/p&gt;
&lt;p&gt;In &lt;em&gt;Half-Life 2&lt;/em&gt;, the Resistance spray-painted λ on walls to mark supply caches you could find later. Your agent needs the same kind of mark: a file on disk that says &amp;ldquo;this iteration already finished,&amp;rdquo; not a memory that evaporates when the turn ends. That is the actual argument for giving an agent a computer instead of a context window.&lt;/p&gt;
&lt;h2 id="how-construct-runs-the-same-job"&gt;How Construct runs the same job&lt;a class="heading-anchor" href="#how-construct-runs-the-same-job" aria-label="Link to this section"&gt;&lt;/a&gt;
&lt;/h2&gt;
&lt;p&gt;Construct&amp;rsquo;s agent has a persistent workspace: a filesystem that outlives the turn, long-term memories it carries across sessions, and reusable workflows that run on demand or from its Calendar. Those three pieces are what turn the client-reporting job from one long bet into a job that converges. Less HEV suit cosplay, more actual hazardous-environment gear: the state that has to survive the run is sitting on disk before anything else starts.&lt;/p&gt;
&lt;p&gt;Concretely, the reporting job becomes one scheduled workflow rather than eight:&lt;/p&gt;
&lt;table&gt;
	&lt;thead&gt;
			&lt;tr&gt;
					&lt;th&gt;Piece&lt;/th&gt;
					&lt;th&gt;What it is in Construct&lt;/th&gt;
					&lt;th&gt;Why it matters to the math&lt;/th&gt;
			&lt;/tr&gt;
	&lt;/thead&gt;
	&lt;tbody&gt;
			&lt;tr&gt;
					&lt;td&gt;The client list&lt;/td&gt;
					&lt;td&gt;A file in the workspace, &lt;code&gt;clients.md&lt;/code&gt;, that the agent reads at the start of every run&lt;/td&gt;
					&lt;td&gt;The run&amp;rsquo;s scope is data on disk, not something rebuilt from a prompt each time&lt;/td&gt;
			&lt;/tr&gt;
			&lt;tr&gt;
					&lt;td&gt;The completed work&lt;/td&gt;
					&lt;td&gt;One file per client under &lt;code&gt;reports/2026-08/&lt;/code&gt;, written as each report finishes&lt;/td&gt;
					&lt;td&gt;Every finished iteration is a checkpoint that outlives the run that made it&lt;/td&gt;
			&lt;/tr&gt;
			&lt;tr&gt;
					&lt;td&gt;The instruction&lt;/td&gt;
					&lt;td&gt;&amp;ldquo;For each client in &lt;code&gt;clients.md&lt;/code&gt; with no report in &lt;code&gt;reports/2026-08/&lt;/code&gt;, produce one and email it&amp;rdquo;&lt;/td&gt;
					&lt;td&gt;The job describes remaining work, so a rerun is a retry rather than a duplicate&lt;/td&gt;
			&lt;/tr&gt;
			&lt;tr&gt;
					&lt;td&gt;Standing preferences&lt;/td&gt;
					&lt;td&gt;Long-term memory: the template a given client wants, which metrics they care about&lt;/td&gt;
					&lt;td&gt;Iteration six does not re-derive what iteration one already established&lt;/td&gt;
			&lt;/tr&gt;
			&lt;tr&gt;
					&lt;td&gt;The retry&lt;/td&gt;
					&lt;td&gt;The schedule itself, plus an on-demand run when you want it sooner&lt;/td&gt;
					&lt;td&gt;The two clients that failed get picked up next pass with no manual bookkeeping&lt;/td&gt;
			&lt;/tr&gt;
			&lt;tr&gt;
					&lt;td&gt;The diagnosis&lt;/td&gt;
					&lt;td&gt;The Activity feed&amp;rsquo;s bounded action summaries and best-effort reasons, plus chat&amp;rsquo;s bounded tool records&lt;/td&gt;
					&lt;td&gt;You find the step that broke instead of rereading a transcript&lt;/td&gt;
			&lt;/tr&gt;
	&lt;/tbody&gt;
&lt;/table&gt;
&lt;p&gt;The important line is the third one. Because the instruction is scoped to clients without a report on disk, a run that dies at client six is not a disaster that needs unwinding. The next run reads the directory, sees five reports, and works on the remaining three. Nothing is sent twice, and nobody has to remember where it stopped.&lt;/p&gt;
&lt;p&gt;Two honest boundaries. Construct&amp;rsquo;s workflows are linear today: branching, fan-out, and subworkflows are not supported, so the clients are worked through inside a run rather than dispatched as eight parallel jobs. And the schedule interval is your retry cadence, so a job that must finish within the hour needs an on-demand rerun rather than patience. The infrastructure that makes a per-user filesystem cheap enough to leave lying around between runs is in &lt;a href="https://construct.computer/blog/running-ai-agents-on-cloudflare-not-vms/" data-external&gt;how our agents get real computers we mostly do not pay for&lt;/a&gt;.&lt;/p&gt;
&lt;p&gt;&lt;em&gt;Read next: &lt;a href="https://construct.computer/blog/ai-workflow-automation/" data-external&gt;AI Workflow Automation&lt;/a&gt;&lt;/em&gt;&lt;/p&gt;
&lt;h2 id="a-checklist-for-building-a-resumable-agent-job"&gt;A checklist for building a resumable agent job&lt;a class="heading-anchor" href="#a-checklist-for-building-a-resumable-agent-job" aria-label="Link to this section"&gt;&lt;/a&gt;
&lt;/h2&gt;
&lt;p&gt;The pattern generalizes past client reports. Anything shaped as &amp;ldquo;for each X, do a few things&amp;rdquo; is a candidate.&lt;/p&gt;
&lt;ol&gt;
&lt;li&gt;&lt;strong&gt;Find the loop.&lt;/strong&gt; Almost every recurring ops job has one: per client, per invoice, per candidate, per repo. That loop boundary is your seam. Cut there first, because it is the cut that turns one long bet into many short ones.&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Make each iteration leave an artifact.&lt;/strong&gt; A file, a CRM record, a sent email. Something a later run can check for existence. If an iteration leaves nothing behind, you cannot tell a retry from a duplicate. Leave a λ on the wall.&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Make iterations skip completed work.&lt;/strong&gt; &amp;ldquo;Write the report for any client that does not already have one dated this month&amp;rdquo; is resumable. &amp;ldquo;Write the reports&amp;rdquo; is not. This one line is usually the difference between a job you can schedule and a job you have to babysit.&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Put the ambiguous judgment in its own step.&lt;/strong&gt; Judgment steps have a lower per-step success rate than mechanical ones. Isolating them means a bad judgment call costs you that step, not the run.&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Supervise the first several runs on demand before scheduling.&lt;/strong&gt; You are looking for which step is your weak link, and you will not guess it correctly.&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Check the trail when something fails.&lt;/strong&gt; Construct&amp;rsquo;s Activity feed keeps bounded action summaries and best-effort reasons, and chat retains bounded tool records, so a failed iteration can be traced to a step rather than to a vibe.&lt;/li&gt;
&lt;/ol&gt;
&lt;p&gt;That third point is worth sitting with. Most people write agent instructions as a description of the work. The instructions that survive scheduling are written as a description of the work that remains.&lt;/p&gt;
&lt;h2 id="what-checkpointing-does-not-fix"&gt;What checkpointing does not fix&lt;a class="heading-anchor" href="#what-checkpointing-does-not-fix" aria-label="Link to this section"&gt;&lt;/a&gt;
&lt;/h2&gt;
&lt;p&gt;Checkpointing lowers the cost of failure. It does not lower the failure rate, and there are jobs where that distinction matters. You still have the same λ. You just refuse to stake the whole sample on one continuous exposure.&lt;/p&gt;
&lt;p&gt;If your steps are genuinely order-dependent all the way through, with no natural loop and no safe stopping point, there is no seam to cut and you are back to one long bet. If a step has an irreversible external side effect, a payment, a public post, a message to a customer, then a retry is not free and idempotency has to be designed in rather than assumed. Construct lets you inspect a running task, interrupt it mid-turn, and answer a question the agent raises, but it does not currently insert a mandatory approval gate before every external side effect, so those steps still need real supervision before they run unattended.&lt;/p&gt;
&lt;p&gt;Memory has a boundary too. Construct keeps workspace files and selected long-term memories across sessions, but automatic memory does not currently index uploaded files, private app results, terminal output, or live browser runs. Context that has to carry across a seam should be written to a file or stored explicitly, not assumed to survive because it appeared once in a transcript.&lt;/p&gt;
&lt;p&gt;And if the work is genuinely deterministic, the same trigger and the same transformation every time with no judgment anywhere in the middle, a rule-based automation platform will do it more cheaply and more predictably than any agent. We laid out where that line falls in &lt;a href="https://construct.computer/blog/ai-agent-vs-zapier/" data-external&gt;AI agent vs Zapier automation&lt;/a&gt;.&lt;/p&gt;
&lt;h2 id="ask-this-instead-of-which-model-should-i-use"&gt;Ask this instead of &amp;ldquo;which model should I use&amp;rdquo;&lt;a class="heading-anchor" href="#ask-this-instead-of-which-model-should-i-use" aria-label="Link to this section"&gt;&lt;/a&gt;
&lt;/h2&gt;
&lt;p&gt;&amp;ldquo;Which model should I use&amp;rdquo; is the question everyone asks. It is not usually the binding constraint. A model one tier better moves your per-step reliability a few points and leaves the exponent untouched. Restructuring a 48-step run into eight 6-step runs moves the exponent, which is where the leverage lives.&lt;/p&gt;
&lt;p&gt;Ask instead: when this job fails at step thirty, what is left on disk, and what does the retry have to redo? If the answer is &amp;ldquo;nothing&amp;rdquo; and &amp;ldquo;all of it,&amp;rdquo; the model is not your problem. You do not need the right man in the wrong place. You need a workspace that outlives the turn.&lt;/p&gt;
&lt;p&gt;Start with &lt;a href="https://construct.computer/blog/ai-workflow-automation/" data-external&gt;AI workflow automation&lt;/a&gt; for how reusable steps and schedules fit together, &lt;a href="https://construct.computer/blog/ai-agent-memory/" data-external&gt;AI agent memory&lt;/a&gt; for what carries across runs and how to correct it, and &lt;a href="https://construct.computer/blog/how-to-choose-an-ai-agent-platform-for-your-team/" data-external&gt;how to choose an AI agent platform&lt;/a&gt; if you are evaluating this against other tools. If the underlying idea is new, &lt;a href="https://construct.computer/blog/what-is-an-ai-employee/" data-external&gt;what is an AI employee&lt;/a&gt; is the place to begin.&lt;/p&gt;
&lt;hr&gt;
&lt;p&gt;&lt;em&gt;Originally posted on &lt;a href="https://construct.computer/blog/agent-task-half-life/" data-external&gt;construct.computer&lt;/a&gt;.&lt;/em&gt;&lt;/p&gt;</content:encoded></item><item><title>All our Agents get computers, we pay for almost none</title><link>https://ankush.one/blogs/running-ai-agents-on-cloudflare-not-vms/</link><pubDate>Fri, 31 Jul 2026 00:00:00 +0000</pubDate><author>ankush4singh@gmail.com (Ankush Singh)</author><guid>https://ankush.one/blogs/running-ai-agents-on-cloudflare-not-vms/</guid><category>ai-agents</category><category>cloudflare</category><category>construct</category><description>Every Construct agent gets a real Linux computer. We pay for almost none of them. How we run that on Cloudflare Durable Objects, Sandboxes, and R2.</description><content:encoded>&lt;p&gt;Construct is an AI employee. You hand it a job. It goes and does the job, without you having to babysit it.&lt;/p&gt;
&lt;p&gt;To do that with agents like Hermes or OpenClaw, you spin up a VPS. Real CPU. Real disk. A full Linux machine that stays on between turns. That costs real money, and the moment more people sign up to poke around and never come back, it &lt;strong&gt;murders your infrastructure bill&lt;/strong&gt;.&lt;/p&gt;
&lt;p&gt;We give every Agent that same kind of computer.&lt;/p&gt;
&lt;p&gt;Don&amp;rsquo;t get excited - we are not paying for all of them.&lt;/p&gt;
&lt;p&gt;The computers work. When someone asks Construct to do something, the machine wakes up and does it. When the tab goes quiet, nothing stays on just to cosplay as useful. Idle boxes are a kink we refuse to subsidize.&lt;/p&gt;
&lt;p&gt;We like edging. So we put the whole stack on the edge. Agent loop in a Durable Object. Linux summoned for one tool call, then gone. Disk that is not a disk. The bill only finishes when somebody actually works.&lt;/p&gt;
&lt;p&gt;This post is how we built that. Behaviour, constants, and tradeoffs are production.&lt;/p&gt;
&lt;h2 id="the-bill-that-scaled-with-people-who-never-came-back"&gt;The bill that scaled with people who never came back&lt;a class="heading-anchor" href="#the-bill-that-scaled-with-people-who-never-came-back" aria-label="Link to this section"&gt;&lt;/a&gt;
&lt;/h2&gt;
&lt;p&gt;The default way to run an AI agent is to hand it a Linux box and walk away.&lt;/p&gt;
&lt;p&gt;Not stupid. Agents are trained on bash. Bash wants a filesystem and a process table. The lazy move is a machine that stays up between turns. Lazy is usually correct.&lt;/p&gt;
&lt;p&gt;Until the invoice arrives.&lt;/p&gt;
&lt;p&gt;We know that invoice because we signed it first. Construct&amp;rsquo;s original backend was a Bun and Elysia monolith on a VPS. SQLite on local disk. nginx out front. Roughly 1,300 lines of container management whose only job was handing every user their own box. It worked. Growing the product meant growing a fleet of real machines with real disks, billed through the silent hours where nobody typed a character.&lt;/p&gt;
&lt;p&gt;That is the expensive version of &amp;ldquo;every user gets a computer.&amp;rdquo; Technically true. Financially a treadmill. Someone who signed up once and ghosted cost the same as a power user. The product could not scale if the bill scaled with accounts instead of work.&lt;/p&gt;
&lt;p&gt;We killed it on 30 March 2026. Commit title: &amp;ldquo;migrate backend from VPS monolith to Cloudflare Workers.&amp;rdquo; Everything in this post starts there.&lt;/p&gt;
&lt;p&gt;Hono on Workers. Per-user WebSocket hub in a Durable Object with hibernation. Container object that archived to R2 on sleep and restored on wake. D1 for relational state. R2 for files. Frontend on Workers assets. Two days later we dragged the leftover container infrastructure into a legacy folder and wrote the plan we still run: the agent runs headlessly inside Durable Objects.&lt;/p&gt;
&lt;p&gt;Containers and Sandboxes went GA on 13 April 2026. We moved onto the Sandbox SDK once it had the shape we wanted.&lt;/p&gt;
&lt;p&gt;We did not see the future. We stared at a cost curve that scaled with signups instead of usage, bet on Durable Objects before there was a tidy product name for the thing, and watched the platform walk to the same answer. camelAI &lt;a href="https://camelai.com/blog/our-coding-agent-runs-in-a-cloudflare-durable-object-not-a-vm" data-external&gt;hit the same wall and reached nearly the same conclusion&lt;/a&gt; independently. Validation, or a warning about people who write posts like this. Both.&lt;/p&gt;
&lt;h2 id="only-half-a-computer-deserves-to-be-awake"&gt;Only half a computer deserves to be awake&lt;a class="heading-anchor" href="#only-half-a-computer-deserves-to-be-awake" aria-label="Link to this section"&gt;&lt;/a&gt;
&lt;/h2&gt;
&lt;p&gt;&amp;ldquo;The agent needs a computer&amp;rdquo; is two unrelated things taped together.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Agent loop:&lt;/strong&gt; transcript, tool routing, model calls, memory recall. This is how Construct thinks across a turn.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Machine:&lt;/strong&gt; the thing that runs &lt;code&gt;pdftotext&lt;/code&gt;, converts a spreadsheet, compiles a project, unzips the archive somebody should not have uploaded. This is how Construct touches a real filesystem.&lt;/p&gt;
&lt;p&gt;Only the second needs Linux.&lt;/p&gt;
&lt;p&gt;The first needs durable state and a socket. Cloudflare sells exactly that. Split the two and the expensive half stops needing to be alive at all.&lt;/p&gt;
&lt;p&gt;Everybody argues about which model to use. Almost nobody argues about which half of their stack is getting paid to nap. That is the whole post, in one sentence.&lt;/p&gt;
&lt;h2 id="linux-shows-up-edges-the-job-and-leaves"&gt;Linux shows up, edges the job, and leaves&lt;a class="heading-anchor" href="#linux-shows-up-edges-the-job-and-leaves" aria-label="Link to this section"&gt;&lt;/a&gt;
&lt;/h2&gt;
&lt;p&gt;This is the section the title is about.&lt;/p&gt;
&lt;p&gt;Each Construct agent gets one Durable Object, resolved by name. Sessions are rows inside it, not objects of their own. Temporary subagents run as delegated sessions on the same parent. Fan out a job - buy concurrency, not instances. That is how Construct can split a big ask without spinning up a fleet.&lt;/p&gt;
&lt;p&gt;The transcript lives in that object&amp;rsquo;s SQLite. Two things, kept apart on purpose. One table is the append-only message log the user scrolls through. A separate record holds the compaction summary plus a watermark of how far it has already consumed. That second one is what reaches the model. Conflate them and you pay, per turn, to re-send an afternoon of conversation to a model that did not ask for it. People do this. At scale.&lt;/p&gt;
&lt;p&gt;Hibernation is where the bill actually moves. We use the WebSocket hibernation API, stash per-connection state on the socket itself, and let the runtime answer keepalives without ever waking the object.&lt;/p&gt;
&lt;p&gt;A tab open all afternoon is not a running process. It is a socket the runtime holds and a handful of SQLite rows. The object wakes when something happens. Goes back down when it does not. Held, not running. Your Construct session feels continuous. Our invoice does not.&lt;/p&gt;
&lt;p&gt;Eviction can happen any moment, so every transcript row also archives to D1. Empty on start? Restore from the archive. Make death cheap and constant and it stops being an incident.&lt;/p&gt;
&lt;p&gt;One annoying edge case: replayed reconnect events can include a turn-started with no matching completion, because the object got evicted mid-turn. Client reconnects, nothing is running - we send terminal idle status so nobody sits watching a spinner think about a thought that died twenty minutes ago.&lt;/p&gt;
&lt;p&gt;When the agent calls the terminal tool, we resolve a sandbox:&lt;/p&gt;
&lt;pre tabindex="0"&gt;&lt;code&gt;sandbox = getSandbox(SANDBOX_BINDING, instanceKeyFor(user, workspace), {
 transport: rpc,
 enableDefaultSession: false, # every exec is a fresh shell
 sleepAfter: idleWindowFor(plan),
})
&lt;/code&gt;&lt;/pre&gt;&lt;p&gt;Instance key scoped to user, or workspace when we know it. Not per agent. Not per session. Live containers track active humans, not agent count. Matters the moment somebody spawns twelve agents and goes to lunch.&lt;/p&gt;
&lt;p&gt;&lt;code&gt;sleepAfter&lt;/code&gt; is the pivot. Ten minutes idle and the box is gone. We call that a feature. You can call it edging. Same config value. The idle window is the plan&amp;rsquo;s command ceiling in minutes, clamped by a global 300-second cap, plus a five minute buffer. That resolves to a ten minute sleep on every plan today. Pair with active CPU billing on Cloudflare Containers - charged for CPU consumed, not wall-clock uptime - and an idle sandbox stops being something you manage.&lt;/p&gt;
&lt;p&gt;Difference between &amp;ldquo;we should build a reaper for idle machines&amp;rdquo; and &amp;ldquo;there is nothing to reap.&amp;rdquo; One is a roadmap item forever. The other is a config value.&lt;/p&gt;
&lt;p&gt;Unglamorous consequence: every exec is a fresh shell. Working directory does not persist. Agents hate this personally. We wrote a hint into the failure path explaining that each terminal call starts clean, relative paths resolve against the scratch tree, and any directory change or exported variable from an earlier call is gone forever. Grief counselling in a string in an error branch.&lt;/p&gt;
&lt;p&gt;Before each exec: kill leftover processes, then mount, then run. Kill after mount and you drop the local bucket&amp;rsquo;s inotify watch and can leave a stale read-only workspace behind. Learned that the hard way.&lt;/p&gt;
&lt;h2 id="she-wrangler-on-my-d1-till-i-r2"&gt;She wrangler on my D1 till I R2&lt;a class="heading-anchor" href="#she-wrangler-on-my-d1-till-i-r2" aria-label="Link to this section"&gt;&lt;/a&gt;
&lt;/h2&gt;
&lt;p&gt;Construct&amp;rsquo;s durable workspace is an R2 bucket mounted into the container over s3fs. &lt;code&gt;allow_other&lt;/code&gt;. Explicit uid and gid so the non-root user can actually write.&lt;/p&gt;
&lt;p&gt;She wrangler on my D1 till I R2. The files finish. The machine does not have to.&lt;/p&gt;
&lt;p&gt;This is the second half of the cost story, and the half people skip. Quotas of 100 MB, 1 GB, and 3 GB are quotas on bytes stored, not volumes kept spun up. When Construct saves a report, a spreadsheet, or a research brief, those files outlive the machine by construction. They were never on the machine. Nobody&amp;rsquo;s work sits on provisioned block storage waiting for their owner to remember the product exists.&lt;/p&gt;
&lt;p&gt;Contrast again: provisioned volume waiting patiently versus bytes in R2 that do not care if the box is awake.&lt;/p&gt;
&lt;p&gt;How we verify the mount matters more than the mount itself. Checking the directory exists is useless. Passes on a stale read-only mount every time. So we make the non-root user touch a probe file and delete it, with a ten second timeout. Exit code zero means the mount is actually writable.&lt;/p&gt;
&lt;p&gt;Probe fails: unmount, remount, fix ownership, probe again, only then give up. Debounced to once a minute. Isolate-level mount cache key carries a version string we bump when mount policy changes, so a warm isolate can never silently skip a policy change.&lt;/p&gt;
&lt;p&gt;Honest footnote: local development uses plain directory sync, not FUSE. Files land root-owned. We chown on mount. That local/prod divergence cost us an afternoon. It is why we trust a write probe over any status output that claims everything is fine.&lt;/p&gt;
&lt;h2 id="what-living-on-the-edge-actually-costs"&gt;What living on the edge actually costs&lt;a class="heading-anchor" href="#what-living-on-the-edge-actually-costs" aria-label="Link to this section"&gt;&lt;/a&gt;
&lt;/h2&gt;
&lt;p&gt;No sales pitch from here down. Living on the edge has a price. We pay it.&lt;/p&gt;
&lt;p&gt;You cannot set &lt;code&gt;USER&lt;/code&gt; in a Cloudflare Sandbox image. Control plane intercepts HTTPS and needs root. Privilege drop moves to the point of execution. Dom energy stays with the control plane. The agent bottoms out as &lt;code&gt;runuser&lt;/code&gt;. Elevated commands stay root through a wrapper. Everything else is &lt;code&gt;runuser -u &amp;lt;agent user&amp;gt;&lt;/code&gt;.&lt;/p&gt;
&lt;p&gt;No sudo. No passwordless sudoers. Process and FD limits by the wrapper. Elevated access per command or until the container sleeps. Grants live inside the container, not Worker memory. Once the machine is disposable, the container is the correct place for short-lived trust. Permission that dies with the box is permission you never have to remember to revoke.&lt;/p&gt;
&lt;p&gt;Memory and a registry catalog used to be separate Workers behind service bindings. We folded them into one. Never a latency optimisation. Service bindings are already efficient in-account transport. The cost was operational: two more deployments, more local processes, more domains, more generated binding contracts, releases that had to be choreographed. Microservice theatre for 41 KiB. We stopped. Nothing came.&lt;/p&gt;
&lt;p&gt;Memory was about 41 KiB compressed. Registry under 1 KiB. Combined Worker about 1,002 KiB against a 2 MiB gate. 500 ms startup stays a cutover check because Wrangler dry runs do not report startup. We wrote down when to undo it: compressed bundle past 2 MiB, startup past 500 ms, own release cadence needed, or failures widen the API blast radius. Four triggers. Before anyone got attached.&lt;/p&gt;
&lt;p&gt;Web app served by a Worker with static assets rather than Pages, because we needed custom domains to manage their own DNS and Pages OAuth could not.&lt;/p&gt;
&lt;p&gt;Living on the edge means never assuming you have time. Construct&amp;rsquo;s memory recall runs on a budget. Whole recall: 800 ms. Inside that, remote embedding plus Vectorize: 250 ms. Those are configured budgets, not measured percentiles. We do not have deployed latency dashboards yet. Docs say so.&lt;/p&gt;
&lt;p&gt;SQLite full-text search and graph traversal are the floor and always return, local to the Durable Object. Vectorize is best-effort on top. Vector path drags: slightly worse answer, not a hang. On one big warm box you would simply wait. Waiting is exactly what you cannot afford when nothing is warm by default. Warmth is a luxury. Edging is the default.&lt;/p&gt;
&lt;p&gt;Embedding model was the one place with real measurements. 18-query set with close distractors, on Workers AI:&lt;/p&gt;
&lt;table&gt;
	&lt;thead&gt;
			&lt;tr&gt;
					&lt;th&gt;Model&lt;/th&gt;
					&lt;th&gt;recall@1&lt;/th&gt;
					&lt;th&gt;mrr&lt;/th&gt;
			&lt;/tr&gt;
	&lt;/thead&gt;
	&lt;tbody&gt;
			&lt;tr&gt;
					&lt;td&gt;qwen3-embedding-0.6b&lt;/td&gt;
					&lt;td&gt;15/18&lt;/td&gt;
					&lt;td&gt;0.892&lt;/td&gt;
			&lt;/tr&gt;
			&lt;tr&gt;
					&lt;td&gt;bge-m3&lt;/td&gt;
					&lt;td&gt;14/18&lt;/td&gt;
					&lt;td&gt;0.889&lt;/td&gt;
			&lt;/tr&gt;
			&lt;tr&gt;
					&lt;td&gt;bge-large-en-v1.5&lt;/td&gt;
					&lt;td&gt;13/18&lt;/td&gt;
					&lt;td&gt;0.852&lt;/td&gt;
			&lt;/tr&gt;
			&lt;tr&gt;
					&lt;td&gt;embeddinggemma-300m&lt;/td&gt;
					&lt;td&gt;12/18&lt;/td&gt;
					&lt;td&gt;0.824&lt;/td&gt;
			&lt;/tr&gt;
	&lt;/tbody&gt;
&lt;/table&gt;
&lt;p&gt;Qwen3 is trained for asymmetric retrieval. Tell it whether it holds a question or a stored record and ranking tightens. Query-side instruction moved mrr from 0.892 to 0.924 on the same set. Free accuracy for one string. Take it.&lt;/p&gt;
&lt;p&gt;A 0.6b model also keeps this inside budget. Memory pipeline gate: USD 1 per 1,000 turns, excluding the primary chat model. Do not swap embedding models casually. Vectorize indexes are built at those exact dimensions. Different model means recreate and re-embed everything.&lt;/p&gt;
&lt;p&gt;Three scars from the platform:&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;&lt;code&gt;setTimeout&lt;/code&gt; versus &lt;code&gt;setAlarm&lt;/code&gt; has a boundary around ten seconds.&lt;/strong&gt; Discord gateway reconnects with backoff. Anything beyond ten seconds goes on an alarm. Timers do not survive eviction. That object holds one WebSocket for every guild we serve. Split it when Discord reports more than one shard, somewhere around 2,500 guilds. Better to schedule that than discover it at 3am.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;A fallback on the same backend as its primary is not a fallback.&lt;/strong&gt; Same outage, different name. Compaction runs the same Gemma build on Workers AI and AI Studio. Failover costs the route. Every other slot crosses providers.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Alarms hand you a bare Durable Object id.&lt;/strong&gt; &lt;code&gt;ctx.id.name&lt;/code&gt; only exists when you reached the object through a named stub. Alarm-triggered instantiation does not. Ours addressed outbound messages. Queue rejected them. Repair alarm re-queued forever, at cost. Infinite unfinished business. Fix: persist the name on first use. Nothing warns you.&lt;/p&gt;
&lt;p&gt;The container is real. It is not free. Image lands between 1.5 and 1.9 GB, which forces &lt;code&gt;standard-1&lt;/code&gt;, because disk allocation is effectively the image size limit. LibreOffice needs 4 GiB or it gets OOM killed. Production configured for up to 100 concurrent instances. Command execution capped at 300 seconds on every plan today. That is a real product limit this architecture caused, not a feature. s3fs is not real POSIX. It will surprise you on a weekend. Fresh shells irritate agents trained on bash. We pay for that in prompt tokens.&lt;/p&gt;
&lt;p&gt;Flat version: if your workload is one long-lived process per user that truly never idles, none of this helps you. Buy a VM. Simpler. Happier. Nobody will make edging jokes at you.&lt;/p&gt;
&lt;h2 id="what-we-run-today"&gt;What we run today&lt;a class="heading-anchor" href="#what-we-run-today" aria-label="Link to this section"&gt;&lt;/a&gt;
&lt;/h2&gt;
&lt;p&gt;Workers. Durable Objects with SQLite storage. Containers through the Sandbox SDK. D1. R2. Queues with dead letter queues. Vectorize. Workers AI. AI Gateway. Email Routing inbound with the send binding outbound. Rate limiting bindings. Cache API. Turnstile. Cron triggers. Static assets. Workers Logs. Agents SDK.&lt;/p&gt;
&lt;p&gt;No KV. No Hyperdrive. No machines of our own. Just Construct users with computers that mostly sleep.&lt;/p&gt;
&lt;p&gt;&lt;a href="https://construct.computer" data-external&gt;construct.computer&lt;/a&gt; is Cloudflare too: Pages, Functions, D1, Turnstile, and the WAF.&lt;/p&gt;
&lt;p&gt;All our Agents get computers, we pay for almost none.&lt;/p&gt;
&lt;p&gt;We like edging. So we put the whole stack on the edge.&lt;/p&gt;
&lt;p&gt;That is how an AI employee can feel always-on without an always-on bill.&lt;/p&gt;
&lt;p&gt;If you want the product side rather than the plumbing, &lt;a href="https://construct.computer/blog/ai-employee/" data-external&gt;what an AI employee actually does&lt;/a&gt; covers the work it completes, &lt;a href="https://construct.computer/blog/ai-agent-memory/" data-external&gt;inspectable agent memory&lt;/a&gt; covers what it remembers and how you correct it, and &lt;a href="https://construct.computer/blog/construct-vs-diy/" data-external&gt;building your own agent stack&lt;/a&gt; is the honest version of what you would be operating instead.&lt;/p&gt;
&lt;hr&gt;
&lt;p&gt;&lt;em&gt;Originally posted on &lt;a href="https://construct.computer/blog/running-ai-agents-on-cloudflare-not-vms/" data-external&gt;construct.computer&lt;/a&gt;.&lt;/em&gt;&lt;/p&gt;</content:encoded></item><item><title>How to Choose an AI Agent Platform for Your Team</title><link>https://ankush.one/blogs/how-to-choose-an-ai-agent-platform-for-your-team/</link><pubDate>Tue, 28 Jul 2026 00:00:00 +0000</pubDate><author>ankush4singh@gmail.com (Ankush Singh)</author><guid>https://ankush.one/blogs/how-to-choose-an-ai-agent-platform-for-your-team/</guid><category>ai-agents</category><category>construct</category><description>Most teams evaluating AI agent platforms compare feature lists and demo videos, then discover months later that the demo and the production job were not the same thing. Here&amp;rsquo;s a framework to choose wisely.</description><content:encoded>&lt;p&gt;Most teams evaluating an AI agent platform compare feature lists and demo videos, then discover months later that the demo and the production job were not the same thing. This is a vendor-agnostic framework for how to choose an AI agent platform: why pilots stall, six criteria that actually predict whether a platform survives contact with real work, the governance rules now attached to that decision, and a scorecard you can apply to any product on your shortlist.&lt;/p&gt;
&lt;h2 id="why-most-ai-agent-pilots-never-reach-production"&gt;Why most AI agent pilots never reach production&lt;a class="heading-anchor" href="#why-most-ai-agent-pilots-never-reach-production" aria-label="Link to this section"&gt;&lt;/a&gt;
&lt;/h2&gt;
&lt;p&gt;An estimated 88% of AI agent pilots fail to reach production, according to an analysis of enterprise deployments compiled by Digital Applied (&lt;a href="https://www.digitalapplied.com/blog/88-percent-ai-agents-never-reach-production-failure-framework" data-external&gt;Digital Applied, AI agent failure framework&lt;/a&gt;). Gartner forecasts the same trend from the vendor side: more than 40% of agentic AI projects will be canceled by the end of 2027 due to escalating costs, unclear business value, or inadequate risk controls (&lt;a href="https://www.digitalapplied.com/blog/agent-washing-definition-buyers-scorecard-2026" data-external&gt;Gartner, press release, cited via Digital Applied&amp;rsquo;s agent-washing scorecard&lt;/a&gt;). Gartner Senior Director Analyst Anushree Verma named the underlying problem directly: &amp;ldquo;most agentic projects today are early-stage experiments driven by hype,&amp;rdquo; which blinds organizations to the real cost and complexity of deploying agents at scale.&lt;/p&gt;
&lt;p&gt;The failure pattern is not random. Digital Applied&amp;rsquo;s breakdown attributes 34% of failed pilots to scope creep, where an initially bounded automation absorbs new requirements until it becomes an open-ended reasoning system nobody scoped for. Data quality failures account for another 27%, when an agent tested against clean sample data meets production records full of incomplete fields and stale formatting. Security and access-control blockers cause 14% of failures, integration complexity accounts for 9%, and governance gaps, missing ownership, monitoring, or incident response, cause another 5% (&lt;a href="https://www.digitalapplied.com/blog/88-percent-ai-agents-never-reach-production-failure-framework" data-external&gt;Digital Applied, AI agent failure framework&lt;/a&gt;). Scope and data readiness alone explain most of the gap between a working demo and a working system.&lt;/p&gt;
&lt;p&gt;This is also an active buying decision for most teams, not a settled one. Only 17% of organizations had deployed AI agents as of 2026, while more than 60% expect to deploy within two years, per Gartner&amp;rsquo;s CIO survey (&lt;a href="https://www.digitalapplied.com/blog/agent-washing-definition-buyers-scorecard-2026" data-external&gt;Digital Applied, agent-washing scorecard&lt;/a&gt;). The evaluation criteria below are aimed at that gap: the difference between a platform that produces a good pilot and one that keeps producing good results after the third team starts depending on it.&lt;/p&gt;
&lt;h2 id="the-ai-agent-evaluation-checklist-6-criteria-that-separate-a-pilot-from-a-platform"&gt;The AI agent evaluation checklist: 6 criteria that separate a pilot from a platform&lt;a class="heading-anchor" href="#the-ai-agent-evaluation-checklist-6-criteria-that-separate-a-pilot-from-a-platform" aria-label="Link to this section"&gt;&lt;/a&gt;
&lt;/h2&gt;
&lt;p&gt;Feature lists answer &amp;ldquo;can it do the task once, in a demo.&amp;rdquo; These six criteria answer &amp;ldquo;will it still be trustworthy after the fifth team is running unsupervised jobs on it.&amp;rdquo; Apply each one to any platform on your shortlist before you commit engineering time to a pilot.&lt;/p&gt;
&lt;h3 id="tool-orchestration-and-execution-surfaces-browser-terminal-files-connected-apps"&gt;Tool orchestration and execution surfaces (browser, terminal, files, connected apps)&lt;a class="heading-anchor" href="#tool-orchestration-and-execution-surfaces-browser-terminal-files-connected-apps" aria-label="Link to this section"&gt;&lt;/a&gt;
&lt;/h3&gt;
&lt;p&gt;Ask what the agent can actually act on, not just what it can talk about. A platform limited to chat and a handful of first-party plugins can explain a task; a platform with a real browser, a sandboxed terminal, persistent files, and structured connected-app actions can complete one. The distinction matters because real jobs mix surfaces: a research step needs the open web, a records update needs a structured app action rather than screen-scraping, and a file transform needs a terminal with somewhere durable to write the result.&lt;/p&gt;
&lt;h3 id="human-in-the-loop-controls-and-approval-gates"&gt;Human-in-the-loop controls and approval gates&lt;a class="heading-anchor" href="#human-in-the-loop-controls-and-approval-gates" aria-label="Link to this section"&gt;&lt;/a&gt;
&lt;/h3&gt;
&lt;p&gt;The question is not whether a platform claims supervision. It is what specifically happens before an agent takes an external action a person did not directly request: can you interrupt a running task, does the agent stop to ask when something is ambiguous, and is there a mandatory approval step before it emails a customer or posts publicly. A platform that answers &amp;ldquo;yes, generally&amp;rdquo; to all three without naming the mechanism has not actually answered the question.&lt;/p&gt;
&lt;h3 id="audit-trails-and-inspectable-activity-and-what-audit-trail-should-actually-mean"&gt;Audit trails and inspectable activity (and what &amp;ldquo;audit trail&amp;rdquo; should actually mean)&lt;a class="heading-anchor" href="#audit-trails-and-inspectable-activity-and-what-audit-trail-should-actually-mean" aria-label="Link to this section"&gt;&lt;/a&gt;
&lt;/h3&gt;
&lt;p&gt;Vendors use &amp;ldquo;audit trail&amp;rdquo; loosely, and the gap between what it implies and what a product delivers matters once something goes wrong. A forensic audit log is an immutable, complete record of every operation, suitable for a compliance investigation. A bounded activity summary is a readable record of what an agent did and roughly why, suitable for an operator reviewing a day&amp;rsquo;s work. Both are useful. They are not the same claim, and a buyer who needs the former should not accept a vendor&amp;rsquo;s description of the latter at face value.&lt;/p&gt;
&lt;h3 id="connector-coverage-and-custom-integrations-via-mcp"&gt;Connector coverage and custom integrations via MCP&lt;a class="heading-anchor" href="#connector-coverage-and-custom-integrations-via-mcp" aria-label="Link to this section"&gt;&lt;/a&gt;
&lt;/h3&gt;
&lt;p&gt;Off-the-shelf integrations cover the common apps, but almost every real deployment eventually hits an internal system with no public API partnership. The Model Context Protocol (MCP) has become the standard way agent platforms reach those systems: about 41% of surveyed software organizations already report agents connected through MCP servers in limited or broad production, according to Stacklok&amp;rsquo;s 2026 survey as reported by Digital Applied (&lt;a href="https://www.digitalapplied.com/blog/mcp-adoption-statistics-2026-model-context-protocol" data-external&gt;Digital Applied, MCP adoption statistics&lt;/a&gt;). Ask whether a platform supports custom MCP servers rather than only a fixed integration catalog, since that is what determines whether an internal tool can be reached without waiting on the vendor&amp;rsquo;s roadmap.&lt;/p&gt;
&lt;h3 id="persistent-correctable-memory-across-sessions"&gt;Persistent, correctable memory across sessions&lt;a class="heading-anchor" href="#persistent-correctable-memory-across-sessions" aria-label="Link to this section"&gt;&lt;/a&gt;
&lt;/h3&gt;
&lt;p&gt;A long chat history is not memory. The useful question is whether a platform can carry forward a stable fact, such as a preferred report format or a client contact, and whether you can correct that fact when it changes without wiping everything else the system knows. A platform that cannot distinguish &amp;ldquo;what was said&amp;rdquo; from &amp;ldquo;what is currently true&amp;rdquo; will eventually act on something that used to be accurate.&lt;/p&gt;
&lt;h3 id="pricing-model-and-real-usage-limits-not-just-a-price-tag"&gt;Pricing model and real usage limits, not just a price tag&lt;a class="heading-anchor" href="#pricing-model-and-real-usage-limits-not-just-a-price-tag" aria-label="Link to this section"&gt;&lt;/a&gt;
&lt;/h3&gt;
&lt;p&gt;A published price answers less than it appears to. The number that actually predicts whether a plan fits a job is what gets metered: steps per task, command runtime, concurrent background jobs, and storage. Two platforms at the same monthly price can support very different workloads depending on where those limits sit, and a plan that looks generous on paper can throttle a deep research task or a long file transform before it finishes.&lt;/p&gt;
&lt;h2 id="governance-pressure-is-raising-the-bar-eu-ai-act-and-agent-washing"&gt;Governance pressure is raising the bar: EU AI Act and &amp;ldquo;agent washing&amp;rdquo;&lt;a class="heading-anchor" href="#governance-pressure-is-raising-the-bar-eu-ai-act-and-agent-washing" aria-label="Link to this section"&gt;&lt;/a&gt;
&lt;/h2&gt;
&lt;p&gt;Two outside pressures are pushing evaluation criteria from nice-to-have to load-bearing. The first is regulatory. The EU AI Act&amp;rsquo;s Article 14 requires high-risk AI systems to be designed so a human can effectively oversee them: monitor operation, intervene, and halt the system through a stop mechanism or equivalent procedure (&lt;a href="https://artificialintelligenceact.eu/article/14/" data-external&gt;EU AI Act, Article 14&lt;/a&gt;). The Act&amp;rsquo;s full set of high-risk obligations, risk management, data governance, logging, transparency, human oversight, and cybersecurity, along with deployer obligations, becomes enforceable on August 2, 2026 (&lt;a href="https://www.legalnodes.com/article/eu-ai-act-2026-updates-compliance-requirements-and-business-risks" data-external&gt;Legal Nodes, EU AI Act 2026 compliance timeline&lt;/a&gt;). A team whose agent work touches a use case the Act classifies as high-risk needs to treat human-in-the-loop controls and inspectable activity as compliance requirements, not product preferences.&lt;/p&gt;
&lt;p&gt;The second pressure is a credibility problem inside the market itself. Gartner estimates that only about 130 of the thousands of vendors marketing &amp;ldquo;agentic AI&amp;rdquo; are genuinely agentic; the rest is what Gartner calls agent washing, existing chatbots, RPA, or assistants relabeled for a hotter category (&lt;a href="https://www.digitalapplied.com/blog/agent-washing-definition-buyers-scorecard-2026" data-external&gt;Digital Applied, agent-washing definition and buyer&amp;rsquo;s scorecard&lt;/a&gt;). That is exactly why the checklist above asks what a platform actually does rather than what it calls itself: a chatbot with a new label still cannot orchestrate tools, hold correctable memory, or leave an inspectable trail, no matter what the pitch deck says.&lt;/p&gt;
&lt;h2 id="questions-to-ask-a-vendor-before-you-pilot"&gt;Questions to ask a vendor before you pilot&lt;a class="heading-anchor" href="#questions-to-ask-a-vendor-before-you-pilot" aria-label="Link to this section"&gt;&lt;/a&gt;
&lt;/h2&gt;
&lt;p&gt;Bring these questions into a vendor call before you commit engineering time to a pilot. Each maps to one of the six criteria above.&lt;/p&gt;
&lt;ol&gt;
&lt;li&gt;&lt;strong&gt;What can the agent act on beyond chat, and what happens to that work when the session ends?&lt;/strong&gt;&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;What specifically stops the agent from doing something it should not?&lt;/strong&gt; Look for an explicit allowlist of capabilities, connected apps, and network origins with runtime denial of calls outside that allowlist.&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Is &amp;ldquo;audit trail&amp;rdquo; a forensic log or a readable summary?&lt;/strong&gt; Get an explicit statement rather than implied feature language.&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Can we reach a system you do not already integrate with?&lt;/strong&gt; Support for custom MCP servers means an internal system can be reached without waiting for first-party support.&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Can the agent be wrong about something and be corrected without losing everything else it knows?&lt;/strong&gt; Look for versioned memory assertions where a correction supersedes a stale fact while keeping history.&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;What is metered, beyond the sticker price?&lt;/strong&gt; Ask for steps per task, command runtime, concurrent jobs, and storage limits stated plainly.&lt;/li&gt;
&lt;/ol&gt;
&lt;h2 id="a-6-point-scorecard-you-can-apply-to-any-ai-agent-platform"&gt;A 6-point scorecard you can apply to any AI agent platform&lt;a class="heading-anchor" href="#a-6-point-scorecard-you-can-apply-to-any-ai-agent-platform" aria-label="Link to this section"&gt;&lt;/a&gt;
&lt;/h2&gt;
&lt;p&gt;Score each criterion against a specific platform before a pilot, not after. A platform that cannot answer a row concretely, only in marketing language, should lose points on that row regardless of how polished the demo looked.&lt;/p&gt;
&lt;table&gt;
	&lt;thead&gt;
			&lt;tr&gt;
					&lt;th&gt;#&lt;/th&gt;
					&lt;th&gt;Criterion&lt;/th&gt;
					&lt;th&gt;Weak signal&lt;/th&gt;
					&lt;th&gt;Strong signal&lt;/th&gt;
			&lt;/tr&gt;
	&lt;/thead&gt;
	&lt;tbody&gt;
			&lt;tr&gt;
					&lt;td&gt;1&lt;/td&gt;
					&lt;td&gt;Execution surfaces&lt;/td&gt;
					&lt;td&gt;Chat plus a fixed plugin list&lt;/td&gt;
					&lt;td&gt;Browser, terminal, files, and connected-app actions in one task&lt;/td&gt;
			&lt;/tr&gt;
			&lt;tr&gt;
					&lt;td&gt;2&lt;/td&gt;
					&lt;td&gt;Human-in-the-loop controls&lt;/td&gt;
					&lt;td&gt;&amp;ldquo;Supervised&amp;rdquo; with no named mechanism&lt;/td&gt;
					&lt;td&gt;Named interrupt, ask-for-input, and approval behavior, with stated gaps&lt;/td&gt;
			&lt;/tr&gt;
			&lt;tr&gt;
					&lt;td&gt;3&lt;/td&gt;
					&lt;td&gt;Audit trail&lt;/td&gt;
					&lt;td&gt;Vague &amp;ldquo;full audit trail&amp;rdquo; claim&lt;/td&gt;
					&lt;td&gt;Explicit statement of whether it is forensic or a bounded summary&lt;/td&gt;
			&lt;/tr&gt;
			&lt;tr&gt;
					&lt;td&gt;4&lt;/td&gt;
					&lt;td&gt;Connector coverage&lt;/td&gt;
					&lt;td&gt;Fixed integration catalog only&lt;/td&gt;
					&lt;td&gt;Native catalog plus custom MCP or an equivalent extension path&lt;/td&gt;
			&lt;/tr&gt;
			&lt;tr&gt;
					&lt;td&gt;5&lt;/td&gt;
					&lt;td&gt;Memory&lt;/td&gt;
					&lt;td&gt;Long chat history described as memory&lt;/td&gt;
					&lt;td&gt;Correctable, provenance-tracked assertions separate from raw history&lt;/td&gt;
			&lt;/tr&gt;
			&lt;tr&gt;
					&lt;td&gt;6&lt;/td&gt;
					&lt;td&gt;Pricing and limits&lt;/td&gt;
					&lt;td&gt;Price with no usage detail&lt;/td&gt;
					&lt;td&gt;Price plus steps, runtime, concurrency, and storage limits stated plainly&lt;/td&gt;
			&lt;/tr&gt;
	&lt;/tbody&gt;
&lt;/table&gt;
&lt;p&gt;A team piloting inbound lead qualification, for example, would run this scorecard against each shortlisted platform using the actual job: read a lead email, decide whether it qualifies, update the CRM, and notify the team in Slack. Score row 1 by watching which surface handles the CRM update. Score row 3 by asking, in writing, whether the vendor&amp;rsquo;s activity record would hold up in a compliance review or only in a weekly stand-up. A platform that scores well on four rows and poorly on two has told you exactly where the pilot needs extra supervision, not whether to abandon it.&lt;/p&gt;
&lt;h2 id="related-resources"&gt;Related resources&lt;a class="heading-anchor" href="#related-resources" aria-label="Link to this section"&gt;&lt;/a&gt;
&lt;/h2&gt;
&lt;ul&gt;
&lt;li&gt;&lt;a href="https://artificialintelligenceact.eu/article/14/" data-external&gt;The EU AI Act&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;a href="https://www.digitalapplied.com/blog/88-percent-ai-agents-never-reach-production-failure-framework" data-external&gt;Digital Applied: AI agent failure framework&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;a href="https://www.digitalapplied.com/blog/agent-washing-definition-buyers-scorecard-2026" data-external&gt;Digital Applied: Agent-washing definition and buyer&amp;rsquo;s scorecard&lt;/a&gt;&lt;/li&gt;
&lt;/ul&gt;
&lt;hr&gt;
&lt;p&gt;&lt;em&gt;Originally posted on &lt;a href="https://construct.computer/blog/how-to-choose-an-ai-agent-platform-for-your-team/" data-external&gt;construct.computer&lt;/a&gt;.&lt;/em&gt;&lt;/p&gt;</content:encoded></item><item><title>How Construct Builds Internal Tools in Your Workspace</title><link>https://ankush.one/blogs/how-construct-builds-internal-tools/</link><pubDate>Mon, 27 Jul 2026 00:00:00 +0000</pubDate><author>ankush4singh@gmail.com (Ankush Singh)</author><guid>https://ankush.one/blogs/how-construct-builds-internal-tools/</guid><category>ai-agents</category><category>construct</category><description>Internal tools often begin as a spreadsheet, a checklist, or a prompt someone runs every week. Construct can turn that repeated process into a small private app.</description><content:encoded>&lt;p&gt;Internal tools often begin as a spreadsheet, a checklist, or a prompt someone runs every week. The process works, but the interface does not: people copy values between tabs, remember status in their heads, and repeat the same setup each time. Construct can turn that repeated process into a small private app that lives inside the same workspace as the files and agent doing the work.&lt;/p&gt;
&lt;p&gt;This is not a one-click path to a public SaaS product. A Construct workspace app is a focused internal interface for one workspace. You describe the job, the agent writes the app, checks that it builds, opens it in the desktop, and keeps the source available for inspection and later changes.&lt;/p&gt;
&lt;h2 id="start-with-the-repeated-process"&gt;Start with the repeated process&lt;a class="heading-anchor" href="#start-with-the-repeated-process" aria-label="Link to this section"&gt;&lt;/a&gt;
&lt;/h2&gt;
&lt;p&gt;The best request describes the work rather than prescribing a software architecture. For example:&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;&amp;ldquo;Build a research request tracker with an owner, priority, status, and link to the finished report.&amp;rdquo;&lt;/li&gt;
&lt;li&gt;&amp;ldquo;Turn this recurring launch checklist into an app the team can update.&amp;rdquo;&lt;/li&gt;
&lt;li&gt;&amp;ldquo;Create a dashboard for the JSON reports stored in this workspace.&amp;rdquo;&lt;/li&gt;
&lt;li&gt;&amp;ldquo;Build an intake form for customer-review requests and keep the queue in one place.&amp;rdquo;&lt;/li&gt;
&lt;/ul&gt;
&lt;p&gt;Construct can start from a list, form, dashboard, or blank interface. The agent creates the app package, writes the React and TypeScript source, declares the capabilities it needs, and validates the result. A successful build opens directly in the Construct web desktop instead of sending you to a separate hosting project.&lt;/p&gt;
&lt;p&gt;The request still matters. &amp;ldquo;Build an operations dashboard&amp;rdquo; leaves important choices unresolved. &amp;ldquo;Show open research requests, let an operator assign an owner, and mark the linked report as reviewed&amp;rdquo; gives the agent a bounded workflow and a clear way to tell whether the app is useful.&lt;/p&gt;
&lt;h2 id="the-app-lives-beside-the-work"&gt;The app lives beside the work&lt;a class="heading-anchor" href="#the-app-lives-beside-the-work" aria-label="Link to this section"&gt;&lt;/a&gt;
&lt;/h2&gt;
&lt;p&gt;A workspace app is not a hidden artifact. Its source lives under an &lt;code&gt;Apps/&amp;lt;app-id&amp;gt;/&lt;/code&gt; folder in Files, where it can be inspected and edited. After a successful build, the app appears as a private workspace install that can be reopened from the desktop, Launchpad, or Spotlight.&lt;/p&gt;
&lt;p&gt;The app&amp;rsquo;s durable state is kept separately under &lt;code&gt;AppData/&amp;lt;app-id&amp;gt;/&lt;/code&gt;. That separation is important: changing the interface should not require overwriting the records the app already manages. Rebuilding can update the code while existing app data remains in the workspace and continues to count against the normal workspace storage quota.&lt;/p&gt;
&lt;p&gt;This model also makes the surrounding agent useful after launch. The same workspace contains the source, relevant files, conversations, and diagnostics. If the interface needs another field or a runtime error appears, the next task can start with the actual app rather than a screenshot and a vague bug report.&lt;/p&gt;
&lt;h2 id="what-a-workspace-app-can-do"&gt;What a workspace app can do&lt;a class="heading-anchor" href="#what-a-workspace-app-can-do" aria-label="Link to this section"&gt;&lt;/a&gt;
&lt;/h2&gt;
&lt;p&gt;The current runtime is intentionally narrower than a general web-hosting platform. That keeps an internal tool close to its workspace purpose.&lt;/p&gt;
&lt;table&gt;
	&lt;thead&gt;
			&lt;tr&gt;
					&lt;th&gt;Capability&lt;/th&gt;
					&lt;th&gt;Current workspace-app support&lt;/th&gt;
			&lt;/tr&gt;
	&lt;/thead&gt;
	&lt;tbody&gt;
			&lt;tr&gt;
					&lt;td&gt;React interface&lt;/td&gt;
					&lt;td&gt;Yes&lt;/td&gt;
			&lt;/tr&gt;
			&lt;tr&gt;
					&lt;td&gt;TypeScript and TSX&lt;/td&gt;
					&lt;td&gt;Yes&lt;/td&gt;
			&lt;/tr&gt;
			&lt;tr&gt;
					&lt;td&gt;Read or write workspace files&lt;/td&gt;
					&lt;td&gt;With declared permission&lt;/td&gt;
			&lt;/tr&gt;
			&lt;tr&gt;
					&lt;td&gt;App-scoped JSON state&lt;/td&gt;
					&lt;td&gt;With declared Files permission&lt;/td&gt;
			&lt;/tr&gt;
			&lt;tr&gt;
					&lt;td&gt;Call an allowed connected app&lt;/td&gt;
					&lt;td&gt;With an explicit app allowlist&lt;/td&gt;
			&lt;/tr&gt;
			&lt;tr&gt;
					&lt;td&gt;Make direct network requests&lt;/td&gt;
					&lt;td&gt;Only to declared HTTPS origins&lt;/td&gt;
			&lt;/tr&gt;
			&lt;tr&gt;
					&lt;td&gt;Install arbitrary npm packages&lt;/td&gt;
					&lt;td&gt;No&lt;/td&gt;
			&lt;/tr&gt;
			&lt;tr&gt;
					&lt;td&gt;Run terminal or agent orchestration from the app UI&lt;/td&gt;
					&lt;td&gt;No&lt;/td&gt;
			&lt;/tr&gt;
			&lt;tr&gt;
					&lt;td&gt;Publish automatically to the public app registry&lt;/td&gt;
					&lt;td&gt;No&lt;/td&gt;
			&lt;/tr&gt;
	&lt;/tbody&gt;
&lt;/table&gt;
&lt;p&gt;The app UI receives a small Construct SDK for workspace files, app-scoped storage, notifications, window controls, and explicitly permitted calls. It does not inherit every tool available to the agent that authored it. The agent can use broader execution surfaces while building, but the finished interface runs with its own narrower contract.&lt;/p&gt;
&lt;h2 id="permissions-are-part-of-the-design"&gt;Permissions are part of the design&lt;a class="heading-anchor" href="#permissions-are-part-of-the-design" aria-label="Link to this section"&gt;&lt;/a&gt;
&lt;/h2&gt;
&lt;p&gt;An internal tool should not gain broad access simply because it was generated inside a trusted workspace. A workspace app declares the native capabilities, connected apps, and network origins it intends to use. Calls outside those allowlists are denied at runtime.&lt;/p&gt;
&lt;p&gt;The platform also rechecks current workspace membership and permissions when the app makes a gateway call. File operations remain subject to the user&amp;rsquo;s workspace access, and an app cannot use its file access to rewrite another app&amp;rsquo;s package. Direct network access is limited to exact HTTPS origins rather than wildcard domains.&lt;/p&gt;
&lt;p&gt;For a request tracker that only reads and writes workspace files, that can mean declaring Files access and nothing else. A tool that posts an approved record to a connected service needs that specific app target as well. The useful default is the smallest permission set that completes the job.&lt;/p&gt;
&lt;h2 id="validation-creates-a-build-not-a-guarantee"&gt;Validation creates a build, not a guarantee&lt;a class="heading-anchor" href="#validation-creates-a-build-not-a-guarantee" aria-label="Link to this section"&gt;&lt;/a&gt;
&lt;/h2&gt;
&lt;p&gt;Before launch, Construct checks the package structure and manifest, follows relative imports, enforces package boundaries, compiles the TypeScript and JSX, and reports design findings. A successful validation creates an immutable artifact with a deterministic build ID and opens the app.&lt;/p&gt;
&lt;p&gt;That proves the package can be built under the workspace runtime. It does not prove that every button, data shape, connected service, or edge case behaves correctly. Runtime verification still matters, especially after changing permissions, storage behavior, or an external call.&lt;/p&gt;
&lt;p&gt;Construct captures bounded console, network, gateway, and runtime-error diagnostics from workspace apps. The agent can inspect recent errors, repair the source, validate again, and ask you to verify the changed behavior. The loop is closer to maintaining a small internal product than generating a disposable code snippet.&lt;/p&gt;
&lt;h2 id="draft-changes-do-not-replace-the-last-good-build"&gt;Draft changes do not replace the last good build&lt;a class="heading-anchor" href="#draft-changes-do-not-replace-the-last-good-build" aria-label="Link to this section"&gt;&lt;/a&gt;
&lt;/h2&gt;
&lt;p&gt;The reliability model is deliberately conservative. A successful build becomes the current runtime artifact. Later source edits mark the app as having draft changes, but opening the app continues to run the last successful build until the new source validates.&lt;/p&gt;
&lt;p&gt;If validation fails, the working artifact is not replaced. Cancelling an edit also leaves both the source and prior build in place. The desktop can show whether an app is ready, has draft changes, or has never built successfully.&lt;/p&gt;
&lt;p&gt;This is a last-good-build fallback, not full version history. It protects the working tool while a repair is in progress; it does not promise browsing or restoring every historical build.&lt;/p&gt;
&lt;h2 id="internal-tool-or-workflow"&gt;Internal tool or workflow?&lt;a class="heading-anchor" href="#internal-tool-or-workflow" aria-label="Link to this section"&gt;&lt;/a&gt;
&lt;/h2&gt;
&lt;p&gt;Not every repeated process needs an interface. Use a workspace app when people need to view, enter, filter, or review state. Use a Construct workflow when the main value is running a known sequence of agent steps, connected-app actions, and notifications. Use a scheduled agent job when the process should run in the background without waiting for someone to open a screen.&lt;/p&gt;
&lt;p&gt;The two can support the same operation. A workspace app can provide the queue and review surface, while files hold durable records and a scheduled job prepares new work. The important boundary is that the app UI does not secretly become an unrestricted agent: automation still runs through the execution surfaces and permissions designed for it.&lt;/p&gt;
&lt;h2 id="where-workspace-apps-fit"&gt;Where workspace apps fit&lt;a class="heading-anchor" href="#where-workspace-apps-fit" aria-label="Link to this section"&gt;&lt;/a&gt;
&lt;/h2&gt;
&lt;p&gt;Good fits are narrow tools with a clear operator and a durable workspace context:&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;Intake forms and review queues&lt;/li&gt;
&lt;li&gt;Checklists for repeated operations&lt;/li&gt;
&lt;li&gt;Lightweight trackers over workspace records&lt;/li&gt;
&lt;li&gt;Dashboards for files or app-scoped JSON data&lt;/li&gt;
&lt;li&gt;Focused interfaces over an explicitly allowed connected service&lt;/li&gt;
&lt;li&gt;Small utilities that make an existing workflow easier to run and inspect&lt;/li&gt;
&lt;/ul&gt;
&lt;p&gt;Poor fits include public customer-facing sites, products that require an arbitrary npm ecosystem, complex multi-service systems, or workloads that need independently operated infrastructure and security controls. Public registry apps also follow a separate developer and review workflow; a private workspace build is not automatically promoted to the public app registry.&lt;/p&gt;
&lt;h2 id="build-the-smallest-useful-interface"&gt;Build the smallest useful interface&lt;a class="heading-anchor" href="#build-the-smallest-useful-interface" aria-label="Link to this section"&gt;&lt;/a&gt;
&lt;/h2&gt;
&lt;p&gt;The shortest path is to start with one repeated process and one person who needs to operate it. Name the records the app should manage, the actions the operator should take, the files or services it may access, and what a successful result looks like. Then ask Construct to build the smallest interface that makes that process easier to run and supervise.&lt;/p&gt;
&lt;p&gt;Once the first build opens, use it. The missing field, unclear status, or unnecessary screen becomes obvious faster in a working tool than in a long specification. Construct can revise the source while the previous good build remains available, giving the internal tool room to improve without turning every edit into a fresh software project.&lt;/p&gt;
&lt;h2 id="related-resources"&gt;Related resources&lt;a class="heading-anchor" href="#related-resources" aria-label="Link to this section"&gt;&lt;/a&gt;
&lt;/h2&gt;
&lt;ul&gt;
&lt;li&gt;&lt;a href="https://construct.computer" data-external&gt;Construct homepage&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;a href="https://construct.computer/blog/what-is-an-ai-employee/" data-external&gt;What is an AI employee?&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;a href="https://construct.computer/blog/ai-workflow-automation/" data-external&gt;AI workflow automation&lt;/a&gt;&lt;/li&gt;
&lt;/ul&gt;
&lt;hr&gt;
&lt;p&gt;&lt;em&gt;Originally posted on &lt;a href="https://construct.computer/blog/build-internal-tools-with-construct/" data-external&gt;construct.computer&lt;/a&gt;.&lt;/em&gt;&lt;/p&gt;</content:encoded></item><item><title>How I built a Pokemon TCG scanner that runs 100% in your browser</title><link>https://ankush.one/blogs/pokemon-scanner/</link><pubDate>Sun, 18 Jan 2026 00:00:00 +0000</pubDate><author>ankush4singh@gmail.com (Ankush Singh)</author><guid>https://ankush.one/blogs/pokemon-scanner/</guid><category>machine-learning</category><category>web</category><description>I wanted a way to scan my Pokemon cards instantly without waiting for server uploads. Here is how I built a fully client-side scanner using modern web AI tools.</description><content:encoded>&lt;figure class="doc-figure doc-figure--portrait"&gt;
 &lt;video autoplay loop muted playsinline poster="https://ankush.one/media/mypokemonscanner-poster.webp" width="332" height="670" aria-label="Scanner Demo"&gt;
 &lt;source src="https://ankush.one/media/mypokemonscanner.mp4" type="video/mp4"&gt;
 &lt;/video&gt;
 &lt;figcaption&gt;It's fast. Really fast.&lt;/figcaption&gt;
&lt;/figure&gt;
&lt;p&gt;I have always been a huge fan of Pokemon. I have played almost every game from GBC to HGSS, and recently played a few 3DS games when I got my hands on a &amp;rsquo;new&amp;rsquo; 3DS in Thailand.&lt;/p&gt;
&lt;p&gt;I also started collecting Pokemon cards last year and as of writing this, I have around 350-400 cards. I couldnt look at them just sitting tightly in a tin and a cardbox, so I bought a pack of penny sleeves and a good binder (without the O-rings) and put all the shiny sparkly holo cards with cool looking artwork at the front, following with reverse holows and other normal cards that I like, and yanked the rest of the bulk back into the tin :)&lt;/p&gt;
&lt;p&gt;At this point I have a bunch of cool art cards, in English, Thai, Japanese, Korean and Chinese, which look soo cool and valuable. I then tried a bunch of apps that scan pokemon cards and tell you more about it, but they all had their limitations or were straightup paywalled, that too by a large margin (not even an amount that one just spends without thinking much), so I did what every engineer would do at that point&lt;/p&gt;
&lt;h2 id="i-hacked-my-own-pokemon-card-scanner-insert-maniacal-laughter"&gt;I hacked my own Pokemon card scanner &lt;em&gt;&amp;lt;insert maniacal laughter&amp;gt;&lt;/em&gt;&lt;a class="heading-anchor" href="#i-hacked-my-own-pokemon-card-scanner-insert-maniacal-laughter" aria-label="Link to this section"&gt;&lt;/a&gt;
&lt;/h2&gt;
&lt;p&gt;TLDR; scraped ~20K card data and images from &lt;a href="https://tcgdex.dev" data-external&gt;TCGdex&lt;/a&gt;, generated embeddings using &lt;a href="https://openai.com/index/clip/" data-external&gt;OpenAI CLIP&lt;/a&gt; model, and used &lt;a href="https://opencv.org/" data-external&gt;OpenCV&lt;/a&gt; to very crudely find a rectangle like region in the image, and crop it out to search in the embeddings.&lt;/p&gt;
&lt;p&gt;While the initial setup worked flawlessly on my mac, it was just a hacky taped together REST API that took an image and returned the card name, card id, image url, and had a bunch of issues I would have to deal with if I wanted to launch this as an app.&lt;/p&gt;
&lt;p&gt;The image embeddings took around 7 minutes using &lt;a href="https://developer.apple.com/metal/pytorch/" data-external&gt;MPS&lt;/a&gt; on my Mac, which is fine, because inference is usually much quicker (~800ms at that time on mac), when I uploaded the embeddings to my VPS without a GPU and served the API, it took multiple seconds to identify the card from an image, which was disastrous. Continuing to use the VPS for reverse image search was the worst thing I could have done.&lt;/p&gt;
&lt;p&gt;I decided I will move everything client side and run everything in the browser- card detection, cropping, recognition, data fetching, everything, and started reading more about different CLIP embedding models- SigLIP, OpenCLIP, MobileCLIP, ViTs, etc etc&amp;hellip; and tried every single one of them. I also learnt about Model Quantization, where you reduce floating point accuracy to reduce the size of model (with some loss in accuracy) and make it run faster with less memory and power. Int8 quantization worked good enough for SigLIP.&lt;/p&gt;
&lt;p&gt;I had already setup a workflow in the process, to generate embeddings and run it on the web using TF.js + onnx web runtime. I used SigLIP for a day or two, deployed a PWA and started messing around with it on my phone, and ended up switching it with MobileClip-S2, which had better results and accuracy recognising pokemon cards. There was some hurdles with quantization, Int8 always broke the exported onnx somehow, but FP16 quntization was fine and I ended up using that (whatever works well right?)&lt;/p&gt;
&lt;p&gt;I also realised that all this time I had been using WebGL/wasm for inference on the web, and added WebGPU backend for Tfjs, which immedieately made inference much much faster.&lt;/p&gt;
&lt;blockquote&gt;
&lt;p&gt;&lt;strong&gt;Sidequest: YOLO model to specifically detect pokemon cards&lt;/strong&gt;&lt;/p&gt;
&lt;p&gt;While opencv rectangle countour thingie to crop out cards works okay for still images, it failed drastically for video streams, which I was trying to use, so I ended up training a YOLO11n model on a pokemon cards dataset I found online, and it worked pretty well and really fast, while also being around 5mb for the entire model&lt;/p&gt;

&lt;/blockquote&gt;&lt;p&gt;By now the inference time on desktop averaged around 10ms for detection using yolo11n and 50-80ms for recognition using MobileClip-S2 and searching the embeddings, with similar times on mobile for yolo and usually a few hundred ms for recognition, but never more than a second if you have a good phone (not guaranteeing anything though)&lt;/p&gt;
&lt;h3 id="cool-now-i-have-a-working-scanner-which-is-fast-runs-client-side"&gt;Cool, now I have a working scanner which is fast runs client side&lt;a class="heading-anchor" href="#cool-now-i-have-a-working-scanner-which-is-fast-runs-client-side" aria-label="Link to this section"&gt;&lt;/a&gt;
&lt;/h3&gt;
&lt;p&gt;What next? Data. TCGplayer has an extensive database on English cards, but Japanese and other languages? next to none &lt;img src="https://ankush.one/emojis/harold.jpg" alt="" loading="lazy" decoding="async"
 style="display:inline; padding:0px; margin:0px; border-radius:2px; border:none; max-height:22px; vertical-align:middle" /&gt; So I had to look for other sources with both card data and images and ended up scraping &lt;a href="https://limitlesstcg.com" data-external&gt;limitlesstcg&lt;/a&gt; for both English and Japanese. I also ended up using the &lt;a href="https://www.tcgplayer.com/" data-external&gt;TCGplayer&lt;/a&gt; apis to fetch pricing info for all cards and sets, including variations like card condition- near mint, lightly played, damaged, card type- normal, holo, reverse holo.&lt;/p&gt;
&lt;p&gt;Now I have a codebase which can:&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;download english card data from tcgdex&lt;/li&gt;
&lt;li&gt;download japanese card data from limitlesstcg&lt;/li&gt;
&lt;li&gt;merge both to fill in missing cards and their info&lt;/li&gt;
&lt;li&gt;fetch whatever price info it can find for our merged cards db&lt;/li&gt;
&lt;li&gt;has apis to search &amp;amp; query our data&lt;/li&gt;
&lt;/ul&gt;
&lt;h2 id="package-and-ship"&gt;Package and ship&lt;a class="heading-anchor" href="#package-and-ship" aria-label="Link to this section"&gt;&lt;/a&gt;
&lt;/h2&gt;
&lt;p&gt;Setup a cron on my VPS to update my card data every 3 days, cleaned up the codebase, wrote QOL scripts, built a nice little frontend hosted it on mypokemonscanner.com and we&amp;rsquo;re done building our card scanner (•̀ᴗ•́ )و&lt;/p&gt;
&lt;p&gt;&lt;em&gt;Update: mypokemonscanner.com is no longer online, but the Wayback Machine has an &lt;a href="https://web.archive.org/web/20260227180613/https://mypokemonscanner.com/" data-external&gt;archived snapshot&lt;/a&gt;. The &lt;a href="https://ankush.one/projects/2026.pokemon-scanner/"&gt;project page&lt;/a&gt; has the demo and the tech stack.&lt;/em&gt;&lt;/p&gt;</content:encoded></item><item><title>Building a DOOM-Style 3D Renderer from Scratch in JavaScript</title><link>https://ankush.one/blogs/doom-style-renderer/</link><pubDate>Wed, 24 Dec 2025 00:00:00 +0000</pubDate><author>ankush4singh@gmail.com (Ankush Singh)</author><guid>https://ankush.one/blogs/doom-style-renderer/</guid><category>graphics</category><category>game-dev</category><category>web</category><category>tutorial</category><description>Ever wondered how classic games like DOOM rendered their immersive 3D worlds? This post breaks down the math and logic behind a retro 3D engine.</description><content:encoded>&lt;p&gt;Links: &lt;a href="https://gist.github.com/ankushKun/653b7eeaf6d4279ef97f1ab6f709110a"&gt;Full Source Code&lt;/a&gt;&lt;/p&gt;
&lt;figure class="doc-figure doc-demo"&gt;
 &lt;div class="doc-demo-frame"&gt;
 &lt;iframe src="https://7jiiceni3iqjhaxe44sesz2qjd5q7m5l3kumbu2gzmoerv5ifi3a.arweave.net/-lCBEajaIJOC5OckSWdQSPsPs6vaqMDTRsscSNeoKjY?" title="DOOM-style renderer demo" loading="lazy"&gt;&lt;/iframe&gt;
 &lt;div class="doc-demo-overlay" aria-hidden="true"&gt;Click to play&lt;/div&gt;
 &lt;/div&gt;
 &lt;figcaption&gt;Move with WASD&lt;/figcaption&gt;
&lt;/figure&gt;

&lt;p&gt;Recently I have been adding games to my portfolio and DOOM is one of them. But instead of just embedding the original game, I wanted to understand how it actually worked. So I went down the rabbit hole (a second time) and built a more whacky version of the doom renderer from scratch and in HTML canvas, no extra dependencies.&lt;/p&gt;
&lt;p&gt;Turns out, classic games like &lt;strong&gt;DOOM&lt;/strong&gt; didn&amp;rsquo;t do real 3D at all. They used something called &lt;strong&gt;2.5D rendering&lt;/strong&gt;, basically faking 3D with clever 2D math. And it ran on hardware way worse (100x or more times worse) than what you are reading this on.&lt;/p&gt;
&lt;p&gt;Today I&amp;rsquo;ll break down how I built my own DOOM-style renderer. No fancy tools required, just a code editor and a web browser.&lt;/p&gt;
&lt;h2 id="is-3d-actually-3d"&gt;Is 3D actually 3D?&lt;a class="heading-anchor" href="#is-3d-actually-3d" aria-label="Link to this section"&gt;&lt;/a&gt;
&lt;/h2&gt;
&lt;p&gt;At its core, 3D rendering is basically &lt;strong&gt;&amp;ldquo;If we have a 3D world and a camera, what does the camera see?&amp;rdquo;&lt;/strong&gt;&lt;/p&gt;
&lt;p&gt;This process generally involves defining the world (objects), placing the camera (viewer), projecting those 3D coordinates onto a flat 2D screen, and finally filling the screen with the correct colors (drawing pixels).&lt;/p&gt;
&lt;div class="highlight"&gt;&lt;pre tabindex="0" class="chroma"&gt;&lt;code class="language-text" data-lang="text"&gt;&lt;span class="line"&gt;&lt;span class="cl"&gt; 3D World Camera 2D Screen
&lt;/span&gt;&lt;/span&gt;&lt;span class="line"&gt;&lt;span class="cl"&gt; ┌─────────┐ ┌─────────┐
&lt;/span&gt;&lt;/span&gt;&lt;span class="line"&gt;&lt;span class="cl"&gt; │ ▓▓▓ │ ────────────► │ ▓▓ │
&lt;/span&gt;&lt;/span&gt;&lt;span class="line"&gt;&lt;span class="cl"&gt; │ ▓▓▓ │ (projection) │ ▓▓▓▓ │
&lt;/span&gt;&lt;/span&gt;&lt;span class="line"&gt;&lt;span class="cl"&gt; │ ▓▓ │ │ ▓▓ │
&lt;/span&gt;&lt;/span&gt;&lt;span class="line"&gt;&lt;span class="cl"&gt; └─────────┘ └─────────┘
&lt;/span&gt;&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;&lt;/div&gt;&lt;h2 id="representing-space-with-a-top-down-2d-map"&gt;Representing Space with a top down 2D Map&lt;a class="heading-anchor" href="#representing-space-with-a-top-down-2d-map" aria-label="Link to this section"&gt;&lt;/a&gt;
&lt;/h2&gt;
&lt;p&gt;The simplest way to define a world map is with a &lt;strong&gt;grid map&lt;/strong&gt;. We can represent this as an array where each cell is either a wall (&lt;code&gt;1&lt;/code&gt;) or empty space (&lt;code&gt;0&lt;/code&gt;).&lt;/p&gt;
&lt;aside class="doc-callout doc-callout--tip" aria-label="Future Scope"&gt;
&lt;p class="doc-callout-title"&gt;&lt;span class="doc-callout-icon"&gt;&lt;svg class="ui-glyph ui-glyph-tip" width="1em" height="1em" viewBox="0 0 16 16" fill="none" stroke="currentColor" stroke-width="1.5" stroke-linecap="round" stroke-linejoin="round" aria-hidden="true" focusable="false"&gt;&lt;path d="M5.5 10.5c-1-.9-1.75-2-1.75-3.5a4.25 4.25 0 0 1 8.5 0c0 1.5-.75 2.6-1.75 3.5v1.5h-5zM6.25 14h3.5"/&gt;&lt;/svg&gt;&lt;/span&gt;&lt;span&gt;Future Scope&lt;/span&gt;&lt;/p&gt;
&lt;div class="doc-callout-body"&gt;
&lt;p&gt;we can implement an &lt;code&gt;enum&lt;/code&gt; and use different types of walls or place objects with different number instead of just 0 and 1.&lt;/p&gt;
&lt;/div&gt;
&lt;/aside&gt;&lt;div class="highlight"&gt;&lt;pre tabindex="0" class="chroma"&gt;&lt;code class="language-javascript" data-lang="javascript"&gt;&lt;span class="line"&gt;&lt;span class="cl"&gt;&lt;span class="kr"&gt;const&lt;/span&gt; &lt;span class="nx"&gt;MAP_WIDTH&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="mi"&gt;11&lt;/span&gt; &lt;span class="c1"&gt;// Number of rows
&lt;/span&gt;&lt;/span&gt;&lt;/span&gt;&lt;span class="line"&gt;&lt;span class="cl"&gt;&lt;span class="kr"&gt;const&lt;/span&gt; &lt;span class="nx"&gt;MAP_HEIGHT&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="mi"&gt;17&lt;/span&gt; &lt;span class="c1"&gt;// Number of columns
&lt;/span&gt;&lt;/span&gt;&lt;/span&gt;&lt;span class="line"&gt;&lt;span class="cl"&gt;&lt;span class="kr"&gt;const&lt;/span&gt; &lt;span class="nx"&gt;TILE_SIZE&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="mi"&gt;25&lt;/span&gt; &lt;span class="c1"&gt;// Size of each tile in world units, increasing this makes the map look more spread out
&lt;/span&gt;&lt;/span&gt;&lt;/span&gt;&lt;span class="line"&gt;&lt;span class="cl"&gt;&lt;span class="kr"&gt;const&lt;/span&gt; &lt;span class="nx"&gt;WALL_HEIGHT&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="mi"&gt;40&lt;/span&gt;
&lt;/span&gt;&lt;/span&gt;&lt;span class="line"&gt;&lt;span class="cl"&gt;
&lt;/span&gt;&lt;/span&gt;&lt;span class="line"&gt;&lt;span class="cl"&gt;&lt;span class="kr"&gt;const&lt;/span&gt; &lt;span class="nx"&gt;MAP&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="p"&gt;[&lt;/span&gt;
&lt;/span&gt;&lt;/span&gt;&lt;span class="line"&gt;&lt;span class="cl"&gt; &lt;span class="mi"&gt;1&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="mi"&gt;1&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="mi"&gt;1&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="mi"&gt;1&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="mi"&gt;1&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="mi"&gt;1&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="mi"&gt;1&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="mi"&gt;1&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="mi"&gt;1&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="mi"&gt;1&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="mi"&gt;1&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
&lt;/span&gt;&lt;/span&gt;&lt;span class="line"&gt;&lt;span class="cl"&gt; &lt;span class="mi"&gt;1&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="mi"&gt;0&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="mi"&gt;0&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="mi"&gt;0&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="mi"&gt;0&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="mi"&gt;0&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="mi"&gt;0&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="mi"&gt;0&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="mi"&gt;0&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="mi"&gt;0&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="mi"&gt;1&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
&lt;/span&gt;&lt;/span&gt;&lt;span class="line"&gt;&lt;span class="cl"&gt; &lt;span class="mi"&gt;1&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="mi"&gt;0&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="mi"&gt;0&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="mi"&gt;0&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="mi"&gt;0&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="mi"&gt;0&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="mi"&gt;0&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="mi"&gt;0&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="mi"&gt;0&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="mi"&gt;0&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="mi"&gt;1&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
&lt;/span&gt;&lt;/span&gt;&lt;span class="line"&gt;&lt;span class="cl"&gt; &lt;span class="mi"&gt;1&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="mi"&gt;0&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="mi"&gt;0&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="mi"&gt;1&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="mi"&gt;0&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="mi"&gt;0&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="mi"&gt;0&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="mi"&gt;1&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="mi"&gt;0&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="mi"&gt;0&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="mi"&gt;1&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
&lt;/span&gt;&lt;/span&gt;&lt;span class="line"&gt;&lt;span class="cl"&gt; &lt;span class="mi"&gt;1&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="mi"&gt;0&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="mi"&gt;0&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="mi"&gt;0&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="mi"&gt;0&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="mi"&gt;0&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="mi"&gt;0&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="mi"&gt;0&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="mi"&gt;0&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="mi"&gt;0&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="mi"&gt;1&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
&lt;/span&gt;&lt;/span&gt;&lt;span class="line"&gt;&lt;span class="cl"&gt; &lt;span class="c1"&gt;// ... more rows
&lt;/span&gt;&lt;/span&gt;&lt;/span&gt;&lt;span class="line"&gt;&lt;span class="cl"&gt;&lt;span class="p"&gt;]&lt;/span&gt;
&lt;/span&gt;&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;&lt;/div&gt;&lt;p&gt;This is effectively a &lt;strong&gt;top-down view&lt;/strong&gt; of the world.&lt;/p&gt;
&lt;div class="highlight"&gt;&lt;pre tabindex="0" class="chroma"&gt;&lt;code class="language-text" data-lang="text"&gt;&lt;span class="line"&gt;&lt;span class="cl"&gt;■■■■■■■■■■■
&lt;/span&gt;&lt;/span&gt;&lt;span class="line"&gt;&lt;span class="cl"&gt;■ ■
&lt;/span&gt;&lt;/span&gt;&lt;span class="line"&gt;&lt;span class="cl"&gt;■ ■
&lt;/span&gt;&lt;/span&gt;&lt;span class="line"&gt;&lt;span class="cl"&gt;■ ■ ■ ■ ← Interior pillars
&lt;/span&gt;&lt;/span&gt;&lt;span class="line"&gt;&lt;span class="cl"&gt;■ ■
&lt;/span&gt;&lt;/span&gt;&lt;span class="line"&gt;&lt;span class="cl"&gt;■ ■
&lt;/span&gt;&lt;/span&gt;&lt;span class="line"&gt;&lt;span class="cl"&gt;■■■■■■■■■■■
&lt;/span&gt;&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;&lt;/div&gt;&lt;h3 id="why-2d-maps-for-3d"&gt;Why 2D Maps for 3D?&lt;a class="heading-anchor" href="#why-2d-maps-for-3d" aria-label="Link to this section"&gt;&lt;/a&gt;
&lt;/h3&gt;
&lt;p&gt;This is one of the techniques games like DOOM used to achieve their 3D effect. The world is fundamentally &lt;strong&gt;2D with height&lt;/strong&gt;. The X and Y coordinates define horizontal position, while height is just a property of the walls rather than a full 3D coordinate.&lt;/p&gt;
&lt;h2 id="player-state"&gt;Player State&lt;a class="heading-anchor" href="#player-state" aria-label="Link to this section"&gt;&lt;/a&gt;
&lt;/h2&gt;
&lt;p&gt;The player (our camera) is defined by three simple values: an &lt;code&gt;x&lt;/code&gt; and &lt;code&gt;y&lt;/code&gt; position in the world, and an &lt;code&gt;angle&lt;/code&gt; representing the direction they are facing.&lt;/p&gt;
&lt;div class="highlight"&gt;&lt;pre tabindex="0" class="chroma"&gt;&lt;code class="language-javascript" data-lang="javascript"&gt;&lt;span class="line"&gt;&lt;span class="cl"&gt;&lt;span class="kr"&gt;const&lt;/span&gt; &lt;span class="nx"&gt;Player&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt;
&lt;/span&gt;&lt;/span&gt;&lt;span class="line"&gt;&lt;span class="cl"&gt; &lt;span class="nx"&gt;x&lt;/span&gt;&lt;span class="o"&gt;:&lt;/span&gt; &lt;span class="mi"&gt;50&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="c1"&gt;// World X position
&lt;/span&gt;&lt;/span&gt;&lt;/span&gt;&lt;span class="line"&gt;&lt;span class="cl"&gt; &lt;span class="nx"&gt;y&lt;/span&gt;&lt;span class="o"&gt;:&lt;/span&gt; &lt;span class="mi"&gt;50&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="c1"&gt;// World Y position 
&lt;/span&gt;&lt;/span&gt;&lt;/span&gt;&lt;span class="line"&gt;&lt;span class="cl"&gt; &lt;span class="nx"&gt;angle&lt;/span&gt;&lt;span class="o"&gt;:&lt;/span&gt; &lt;span class="mi"&gt;0&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="c1"&gt;// Direction facing
&lt;/span&gt;&lt;/span&gt;&lt;/span&gt;&lt;span class="line"&gt;&lt;span class="cl"&gt;&lt;span class="p"&gt;}&lt;/span&gt;
&lt;/span&gt;&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;&lt;/div&gt;&lt;h3 id="coordinate-system"&gt;Coordinate System&lt;a class="heading-anchor" href="#coordinate-system" aria-label="Link to this section"&gt;&lt;/a&gt;
&lt;/h3&gt;
&lt;p&gt;We use a standard coordinate system where angle 0° looks toward positive Y (down on the map) and 90° looks toward positive X (right on the map). All angles are in degrees (0-359).&lt;/p&gt;
&lt;div class="highlight"&gt;&lt;pre tabindex="0" class="chroma"&gt;&lt;code class="language-text" data-lang="text"&gt;&lt;span class="line"&gt;&lt;span class="cl"&gt; 180° (0,0) ────────► X
&lt;/span&gt;&lt;/span&gt;&lt;span class="line"&gt;&lt;span class="cl"&gt; ▲ │
&lt;/span&gt;&lt;/span&gt;&lt;span class="line"&gt;&lt;span class="cl"&gt; │ │
&lt;/span&gt;&lt;/span&gt;&lt;span class="line"&gt;&lt;span class="cl"&gt; │ │
&lt;/span&gt;&lt;/span&gt;&lt;span class="line"&gt;&lt;span class="cl"&gt;270° ◄────┼────► 90° │
&lt;/span&gt;&lt;/span&gt;&lt;span class="line"&gt;&lt;span class="cl"&gt; │ ▼
&lt;/span&gt;&lt;/span&gt;&lt;span class="line"&gt;&lt;span class="cl"&gt; ▼ Y
&lt;/span&gt;&lt;/span&gt;&lt;span class="line"&gt;&lt;span class="cl"&gt; 0°
&lt;/span&gt;&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;&lt;/div&gt;&lt;p&gt;For performance, we pre-calculate degree-to-radian conversions so we don&amp;rsquo;t have to constantly do the math during the render loop.&lt;/p&gt;
&lt;div class="highlight"&gt;&lt;pre tabindex="0" class="chroma"&gt;&lt;code class="language-javascript" data-lang="javascript"&gt;&lt;span class="line"&gt;&lt;span class="cl"&gt;&lt;span class="kr"&gt;const&lt;/span&gt; &lt;span class="nx"&gt;RAD&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="nb"&gt;Math&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nx"&gt;PI&lt;/span&gt; &lt;span class="o"&gt;/&lt;/span&gt; &lt;span class="mi"&gt;180&lt;/span&gt;
&lt;/span&gt;&lt;/span&gt;&lt;span class="line"&gt;&lt;span class="cl"&gt;&lt;span class="kr"&gt;const&lt;/span&gt; &lt;span class="nx"&gt;DEG_TO_RAD&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="nb"&gt;Array&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nx"&gt;from&lt;/span&gt;&lt;span class="p"&gt;({&lt;/span&gt; &lt;span class="nx"&gt;length&lt;/span&gt;&lt;span class="o"&gt;:&lt;/span&gt; &lt;span class="mi"&gt;360&lt;/span&gt; &lt;span class="p"&gt;},&lt;/span&gt; &lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="nx"&gt;_&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="nx"&gt;i&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt; &lt;span class="p"&gt;=&amp;gt;&lt;/span&gt; &lt;span class="nx"&gt;i&lt;/span&gt; &lt;span class="o"&gt;*&lt;/span&gt; &lt;span class="nx"&gt;RAD&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
&lt;/span&gt;&lt;/span&gt;&lt;span class="line"&gt;&lt;span class="cl"&gt;
&lt;/span&gt;&lt;/span&gt;&lt;span class="line"&gt;&lt;span class="cl"&gt;&lt;span class="c1"&gt;// Usage: instead of calculating Math.sin(angle * Math.PI / 180)
&lt;/span&gt;&lt;/span&gt;&lt;/span&gt;&lt;span class="line"&gt;&lt;span class="cl"&gt;&lt;span class="c1"&gt;// We just do: Math.sin(DEG_TO_RAD[angle])
&lt;/span&gt;&lt;/span&gt;&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;&lt;/div&gt;&lt;h2 id="transforming-world-to-screen"&gt;Transforming World to Screen&lt;a class="heading-anchor" href="#transforming-world-to-screen" aria-label="Link to this section"&gt;&lt;/a&gt;
&lt;/h2&gt;
&lt;p&gt;This is where the 3D illusion happens! We need to transform world coordinates to screen coordinates in two steps.&lt;/p&gt;
&lt;h3 id="step-1-translate-move-world-relative-to-player"&gt;Step 1: Translate (Move world relative to player)&lt;a class="heading-anchor" href="#step-1-translate-move-world-relative-to-player" aria-label="Link to this section"&gt;&lt;/a&gt;
&lt;/h3&gt;
&lt;p&gt;First, we move everything so the player is at the origin &lt;code&gt;(0, 0)&lt;/code&gt;. We simply subtract the player&amp;rsquo;s position from the wall&amp;rsquo;s position.&lt;/p&gt;
&lt;p&gt;So instead of actually moving the player, we are moving the world relative to the player. From the players perspective, it seems they are moving.&lt;/p&gt;
&lt;div class="highlight"&gt;&lt;pre tabindex="0" class="chroma"&gt;&lt;code class="language-javascript" data-lang="javascript"&gt;&lt;span class="line"&gt;&lt;span class="cl"&gt;&lt;span class="kr"&gt;const&lt;/span&gt; &lt;span class="nx"&gt;relX&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="nx"&gt;wallX&lt;/span&gt; &lt;span class="o"&gt;-&lt;/span&gt; &lt;span class="nx"&gt;Player&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nx"&gt;x&lt;/span&gt;
&lt;/span&gt;&lt;/span&gt;&lt;span class="line"&gt;&lt;span class="cl"&gt;&lt;span class="kr"&gt;const&lt;/span&gt; &lt;span class="nx"&gt;relY&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="nx"&gt;wallY&lt;/span&gt; &lt;span class="o"&gt;-&lt;/span&gt; &lt;span class="nx"&gt;Player&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nx"&gt;y&lt;/span&gt;
&lt;/span&gt;&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;&lt;/div&gt;&lt;h3 id="step-2-rotate-align-world-objects-with-viewing-angle"&gt;Step 2: Rotate (Align world objects with viewing angle)&lt;a class="heading-anchor" href="#step-2-rotate-align-world-objects-with-viewing-angle" aria-label="Link to this section"&gt;&lt;/a&gt;
&lt;/h3&gt;
&lt;p&gt;We rotate the worlds objects relative to the player&amp;rsquo;s viewing angle.&lt;/p&gt;
&lt;div class="highlight"&gt;&lt;pre tabindex="0" class="chroma"&gt;&lt;code class="language-javascript" data-lang="javascript"&gt;&lt;span class="line"&gt;&lt;span class="cl"&gt;&lt;span class="kr"&gt;const&lt;/span&gt; &lt;span class="nx"&gt;ps&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="nb"&gt;Math&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nx"&gt;sin&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="nx"&gt;DEG_TO_RAD&lt;/span&gt;&lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="nx"&gt;Player&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nx"&gt;angle&lt;/span&gt;&lt;span class="p"&gt;])&lt;/span&gt; &lt;span class="c1"&gt;// sin of player angle
&lt;/span&gt;&lt;/span&gt;&lt;/span&gt;&lt;span class="line"&gt;&lt;span class="cl"&gt;&lt;span class="kr"&gt;const&lt;/span&gt; &lt;span class="nx"&gt;pc&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="nb"&gt;Math&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nx"&gt;cos&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="nx"&gt;DEG_TO_RAD&lt;/span&gt;&lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="nx"&gt;Player&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nx"&gt;angle&lt;/span&gt;&lt;span class="p"&gt;])&lt;/span&gt; &lt;span class="c1"&gt;// cos of player angle
&lt;/span&gt;&lt;/span&gt;&lt;/span&gt;&lt;span class="line"&gt;&lt;span class="cl"&gt;
&lt;/span&gt;&lt;/span&gt;&lt;span class="line"&gt;&lt;span class="cl"&gt;&lt;span class="c1"&gt;// Rotation matrix application
&lt;/span&gt;&lt;/span&gt;&lt;/span&gt;&lt;span class="line"&gt;&lt;span class="cl"&gt;&lt;span class="kr"&gt;const&lt;/span&gt; &lt;span class="nx"&gt;tx&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="nx"&gt;relX&lt;/span&gt; &lt;span class="o"&gt;*&lt;/span&gt; &lt;span class="nx"&gt;pc&lt;/span&gt; &lt;span class="o"&gt;+&lt;/span&gt; &lt;span class="nx"&gt;relY&lt;/span&gt; &lt;span class="o"&gt;*&lt;/span&gt; &lt;span class="nx"&gt;ps&lt;/span&gt; &lt;span class="c1"&gt;// transformed X (left/right)
&lt;/span&gt;&lt;/span&gt;&lt;/span&gt;&lt;span class="line"&gt;&lt;span class="cl"&gt;&lt;span class="kr"&gt;const&lt;/span&gt; &lt;span class="nx"&gt;ty&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="nx"&gt;relY&lt;/span&gt; &lt;span class="o"&gt;*&lt;/span&gt; &lt;span class="nx"&gt;pc&lt;/span&gt; &lt;span class="o"&gt;-&lt;/span&gt; &lt;span class="nx"&gt;relX&lt;/span&gt; &lt;span class="o"&gt;*&lt;/span&gt; &lt;span class="nx"&gt;ps&lt;/span&gt; &lt;span class="c1"&gt;// transformed Y (depth/distance)
&lt;/span&gt;&lt;/span&gt;&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;&lt;/div&gt;&lt;p&gt;&lt;strong&gt;What this means:&lt;/strong&gt; &lt;code&gt;tx&lt;/code&gt; is how far left or right the point is from center of view, and &lt;code&gt;ty&lt;/code&gt; is how far forward the point is (the &lt;strong&gt;depth&lt;/strong&gt;).&lt;/p&gt;
&lt;h2 id="projecting-3d-data-onto-2d-screen"&gt;Projecting 3D data onto 2D screen&lt;a class="heading-anchor" href="#projecting-3d-data-onto-2d-screen" aria-label="Link to this section"&gt;&lt;/a&gt;
&lt;/h2&gt;
&lt;p&gt;The furthest the image is, the smaller it should look. Mathematically, this just means dividing by depth (&lt;code&gt;ty&lt;/code&gt;).&lt;/p&gt;
&lt;div class="highlight"&gt;&lt;pre tabindex="0" class="chroma"&gt;&lt;code class="language-javascript" data-lang="javascript"&gt;&lt;span class="line"&gt;&lt;span class="cl"&gt;&lt;span class="kr"&gt;const&lt;/span&gt; &lt;span class="nx"&gt;fov&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="mi"&gt;90&lt;/span&gt; &lt;span class="c1"&gt;// Field of view scalar
&lt;/span&gt;&lt;/span&gt;&lt;/span&gt;&lt;span class="line"&gt;&lt;span class="cl"&gt;
&lt;/span&gt;&lt;/span&gt;&lt;span class="line"&gt;&lt;span class="cl"&gt;&lt;span class="c1"&gt;// Screen X (horizontal position)
&lt;/span&gt;&lt;/span&gt;&lt;/span&gt;&lt;span class="line"&gt;&lt;span class="cl"&gt;&lt;span class="c1"&gt;// Add screen.width/2 to center it
&lt;/span&gt;&lt;/span&gt;&lt;/span&gt;&lt;span class="line"&gt;&lt;span class="cl"&gt;&lt;span class="kr"&gt;const&lt;/span&gt; &lt;span class="nx"&gt;screenX&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="nx"&gt;tx&lt;/span&gt; &lt;span class="o"&gt;*&lt;/span&gt; &lt;span class="nx"&gt;fov&lt;/span&gt; &lt;span class="o"&gt;/&lt;/span&gt; &lt;span class="nx"&gt;ty&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt; &lt;span class="o"&gt;+&lt;/span&gt; &lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="nx"&gt;screen&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nx"&gt;width&lt;/span&gt; &lt;span class="o"&gt;/&lt;/span&gt; &lt;span class="mi"&gt;2&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
&lt;/span&gt;&lt;/span&gt;&lt;span class="line"&gt;&lt;span class="cl"&gt;
&lt;/span&gt;&lt;/span&gt;&lt;span class="line"&gt;&lt;span class="cl"&gt;&lt;span class="c1"&gt;// Screen Y (height)
&lt;/span&gt;&lt;/span&gt;&lt;/span&gt;&lt;span class="line"&gt;&lt;span class="cl"&gt;&lt;span class="c1"&gt;// Taller walls or closer walls appear larger
&lt;/span&gt;&lt;/span&gt;&lt;/span&gt;&lt;span class="line"&gt;&lt;span class="cl"&gt;&lt;span class="kr"&gt;const&lt;/span&gt; &lt;span class="nx"&gt;screenHeight&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="nx"&gt;WALL_HEIGHT&lt;/span&gt; &lt;span class="o"&gt;*&lt;/span&gt; &lt;span class="nx"&gt;fov&lt;/span&gt; &lt;span class="o"&gt;/&lt;/span&gt; &lt;span class="nx"&gt;ty&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
&lt;/span&gt;&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;&lt;/div&gt;&lt;h2 id="translate-walls"&gt;Translate walls&lt;a class="heading-anchor" href="#translate-walls" aria-label="Link to this section"&gt;&lt;/a&gt;
&lt;/h2&gt;
&lt;p&gt;To draw a wall, we need four points on the screen: Top-Left, Top-Right, Bottom-Left, and Bottom-Right.&lt;/p&gt;
&lt;div class="highlight"&gt;&lt;pre tabindex="0" class="chroma"&gt;&lt;code class="language-javascript" data-lang="javascript"&gt;&lt;span class="line"&gt;&lt;span class="cl"&gt;&lt;span class="kd"&gt;function&lt;/span&gt; &lt;span class="nx"&gt;project&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="nx"&gt;x&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="nx"&gt;y&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="nx"&gt;z&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt;
&lt;/span&gt;&lt;/span&gt;&lt;span class="line"&gt;&lt;span class="cl"&gt; &lt;span class="c1"&gt;// z is height (0 for floor, WALL_HEIGHT for ceiling)
&lt;/span&gt;&lt;/span&gt;&lt;/span&gt;&lt;span class="line"&gt;&lt;span class="cl"&gt; &lt;span class="c1"&gt;// ... translation logic ...
&lt;/span&gt;&lt;/span&gt;&lt;/span&gt;&lt;span class="line"&gt;&lt;span class="cl"&gt; &lt;span class="c1"&gt;// ... rotation logic ...
&lt;/span&gt;&lt;/span&gt;&lt;/span&gt;&lt;span class="line"&gt;&lt;span class="cl"&gt; 
&lt;/span&gt;&lt;/span&gt;&lt;span class="line"&gt;&lt;span class="cl"&gt; &lt;span class="c1"&gt;// Projection
&lt;/span&gt;&lt;/span&gt;&lt;/span&gt;&lt;span class="line"&gt;&lt;span class="cl"&gt; &lt;span class="kr"&gt;const&lt;/span&gt; &lt;span class="nx"&gt;screenX&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="nx"&gt;tx&lt;/span&gt; &lt;span class="o"&gt;*&lt;/span&gt; &lt;span class="nx"&gt;fov&lt;/span&gt; &lt;span class="o"&gt;/&lt;/span&gt; &lt;span class="nx"&gt;ty&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt; &lt;span class="o"&gt;+&lt;/span&gt; &lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="nx"&gt;screen&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nx"&gt;width&lt;/span&gt; &lt;span class="o"&gt;/&lt;/span&gt; &lt;span class="mi"&gt;2&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
&lt;/span&gt;&lt;/span&gt;&lt;span class="line"&gt;&lt;span class="cl"&gt; &lt;span class="kr"&gt;const&lt;/span&gt; &lt;span class="nx"&gt;screenY&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="nx"&gt;z&lt;/span&gt; &lt;span class="o"&gt;*&lt;/span&gt; &lt;span class="nx"&gt;fov&lt;/span&gt; &lt;span class="o"&gt;/&lt;/span&gt; &lt;span class="nx"&gt;ty&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt; &lt;span class="o"&gt;+&lt;/span&gt; &lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="nx"&gt;screen&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nx"&gt;height&lt;/span&gt; &lt;span class="o"&gt;/&lt;/span&gt; &lt;span class="mi"&gt;2&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
&lt;/span&gt;&lt;/span&gt;&lt;span class="line"&gt;&lt;span class="cl"&gt; 
&lt;/span&gt;&lt;/span&gt;&lt;span class="line"&gt;&lt;span class="cl"&gt; &lt;span class="k"&gt;return&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt; &lt;span class="nx"&gt;x&lt;/span&gt;&lt;span class="o"&gt;:&lt;/span&gt; &lt;span class="nx"&gt;screenX&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="nx"&gt;y&lt;/span&gt;&lt;span class="o"&gt;:&lt;/span&gt; &lt;span class="nx"&gt;screenY&lt;/span&gt; &lt;span class="p"&gt;}&lt;/span&gt;
&lt;/span&gt;&lt;/span&gt;&lt;span class="line"&gt;&lt;span class="cl"&gt;&lt;span class="p"&gt;}&lt;/span&gt;
&lt;/span&gt;&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;&lt;/div&gt;&lt;p&gt;We project the two ends of a wall segment to get their screen coordinates, then fill the polygon between them using &lt;code&gt;ctx.beginPath()&lt;/code&gt; and &lt;code&gt;ctx.fill()&lt;/code&gt;.&lt;/p&gt;
&lt;h2 id="sort-the-walls-and-finally-draw-them-on-canvas"&gt;Sort the walls and finally draw them on canvas&lt;a class="heading-anchor" href="#sort-the-walls-and-finally-draw-them-on-canvas" aria-label="Link to this section"&gt;&lt;/a&gt;
&lt;/h2&gt;
&lt;p&gt;If we just draw walls in any order, distant walls might be drawn &lt;em&gt;on top&lt;/em&gt; of close walls. That would look weird, so we first sort the walls based on their distance from the player.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;The Painter&amp;rsquo;s Algorithm:&lt;/strong&gt; Draw background objects first, then foreground objects over them.&lt;/p&gt;
&lt;div class="highlight"&gt;&lt;pre tabindex="0" class="chroma"&gt;&lt;code class="language-javascript" data-lang="javascript"&gt;&lt;span class="line"&gt;&lt;span class="cl"&gt;&lt;span class="kd"&gt;function&lt;/span&gt; &lt;span class="nx"&gt;draw3D&lt;/span&gt;&lt;span class="p"&gt;()&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt;
&lt;/span&gt;&lt;/span&gt;&lt;span class="line"&gt;&lt;span class="cl"&gt; &lt;span class="kr"&gt;const&lt;/span&gt; &lt;span class="nx"&gt;wallsWithDistance&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="nx"&gt;walls&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nx"&gt;map&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="nx"&gt;wall&lt;/span&gt; &lt;span class="p"&gt;=&amp;gt;&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt;
&lt;/span&gt;&lt;/span&gt;&lt;span class="line"&gt;&lt;span class="cl"&gt; &lt;span class="c1"&gt;// Calculate distance to player
&lt;/span&gt;&lt;/span&gt;&lt;/span&gt;&lt;span class="line"&gt;&lt;span class="cl"&gt; &lt;span class="kr"&gt;const&lt;/span&gt; &lt;span class="nx"&gt;dist1&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="nx"&gt;dx1&lt;/span&gt;&lt;span class="o"&gt;*&lt;/span&gt;&lt;span class="nx"&gt;dx1&lt;/span&gt; &lt;span class="o"&gt;+&lt;/span&gt; &lt;span class="nx"&gt;dy1&lt;/span&gt;&lt;span class="o"&gt;*&lt;/span&gt;&lt;span class="nx"&gt;dy1&lt;/span&gt;
&lt;/span&gt;&lt;/span&gt;&lt;span class="line"&gt;&lt;span class="cl"&gt; &lt;span class="kr"&gt;const&lt;/span&gt; &lt;span class="nx"&gt;dist2&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="nx"&gt;dx2&lt;/span&gt;&lt;span class="o"&gt;*&lt;/span&gt;&lt;span class="nx"&gt;dx2&lt;/span&gt; &lt;span class="o"&gt;+&lt;/span&gt; &lt;span class="nx"&gt;dy2&lt;/span&gt;&lt;span class="o"&gt;*&lt;/span&gt;&lt;span class="nx"&gt;dy2&lt;/span&gt;
&lt;/span&gt;&lt;/span&gt;&lt;span class="line"&gt;&lt;span class="cl"&gt; 
&lt;/span&gt;&lt;/span&gt;&lt;span class="line"&gt;&lt;span class="cl"&gt; &lt;span class="c1"&gt;// ... calculate weighted distance ...
&lt;/span&gt;&lt;/span&gt;&lt;/span&gt;&lt;span class="line"&gt;&lt;span class="cl"&gt; &lt;span class="k"&gt;return&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt; &lt;span class="nx"&gt;wall&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="nx"&gt;distance&lt;/span&gt;&lt;span class="o"&gt;:&lt;/span&gt; &lt;span class="nx"&gt;weightedDist&lt;/span&gt; &lt;span class="p"&gt;}&lt;/span&gt;
&lt;/span&gt;&lt;/span&gt;&lt;span class="line"&gt;&lt;span class="cl"&gt; &lt;span class="p"&gt;})&lt;/span&gt;
&lt;/span&gt;&lt;/span&gt;&lt;span class="line"&gt;&lt;span class="cl"&gt; 
&lt;/span&gt;&lt;/span&gt;&lt;span class="line"&gt;&lt;span class="cl"&gt; &lt;span class="c1"&gt;// Sort far to near
&lt;/span&gt;&lt;/span&gt;&lt;/span&gt;&lt;span class="line"&gt;&lt;span class="cl"&gt; &lt;span class="nx"&gt;wallsWithDistance&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nx"&gt;sort&lt;/span&gt;&lt;span class="p"&gt;((&lt;/span&gt;&lt;span class="nx"&gt;a&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="nx"&gt;b&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt; &lt;span class="p"&gt;=&amp;gt;&lt;/span&gt; &lt;span class="nx"&gt;b&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nx"&gt;distance&lt;/span&gt; &lt;span class="o"&gt;-&lt;/span&gt; &lt;span class="nx"&gt;a&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nx"&gt;distance&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
&lt;/span&gt;&lt;/span&gt;&lt;span class="line"&gt;&lt;span class="cl"&gt; 
&lt;/span&gt;&lt;/span&gt;&lt;span class="line"&gt;&lt;span class="cl"&gt; &lt;span class="c1"&gt;// Draw
&lt;/span&gt;&lt;/span&gt;&lt;/span&gt;&lt;span class="line"&gt;&lt;span class="cl"&gt; &lt;span class="nx"&gt;wallsWithDistance&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nx"&gt;forEach&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="nx"&gt;item&lt;/span&gt; &lt;span class="p"&gt;=&amp;gt;&lt;/span&gt; &lt;span class="nx"&gt;drawWall&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="nx"&gt;item&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nx"&gt;wall&lt;/span&gt;&lt;span class="p"&gt;))&lt;/span&gt;
&lt;/span&gt;&lt;/span&gt;&lt;span class="line"&gt;&lt;span class="cl"&gt;&lt;span class="p"&gt;}&lt;/span&gt;
&lt;/span&gt;&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;&lt;/div&gt;&lt;h3 id="why-the-weighted-distance"&gt;Why the Weighted Distance?&lt;a class="heading-anchor" href="#why-the-weighted-distance" aria-label="Link to this section"&gt;&lt;/a&gt;
&lt;/h3&gt;
&lt;p&gt;Simple closest-point distance can fail for angled walls. The weighted average considers the closest point (most important for occlusion), the midpoint, and the farthest point to handle edge cases gracefully.&lt;/p&gt;
&lt;h2 id="clipping-dealing-with-edge-cases"&gt;Clipping: Dealing with Edge Cases&lt;a class="heading-anchor" href="#clipping-dealing-with-edge-cases" aria-label="Link to this section"&gt;&lt;/a&gt;
&lt;/h2&gt;
&lt;h3 id="the-problem-behind-the-camera"&gt;The Problem: Behind the Camera&lt;a class="heading-anchor" href="#the-problem-behind-the-camera" aria-label="Link to this section"&gt;&lt;/a&gt;
&lt;/h3&gt;
&lt;p&gt;When a wall is partially behind the player, the math breaks! Computing &lt;code&gt;screenX = tx * fov / ty&lt;/code&gt; when &lt;code&gt;ty&lt;/code&gt; is negative or zero creates division by zero or negative numbers, resulting in visual glitches.&lt;/p&gt;
&lt;h3 id="near-plane-clipping"&gt;Near-Plane Clipping&lt;a class="heading-anchor" href="#near-plane-clipping" aria-label="Link to this section"&gt;&lt;/a&gt;
&lt;/h3&gt;
&lt;p&gt;We define a &lt;strong&gt;near plane&lt;/strong&gt; &amp;ndash; a minimum distance. Anything closer gets clipped.&lt;/p&gt;
&lt;div class="highlight"&gt;&lt;pre tabindex="0" class="chroma"&gt;&lt;code class="language-javascript" data-lang="javascript"&gt;&lt;span class="line"&gt;&lt;span class="cl"&gt;&lt;span class="kr"&gt;const&lt;/span&gt; &lt;span class="nx"&gt;NEAR_PLANE&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="mi"&gt;1&lt;/span&gt;
&lt;/span&gt;&lt;/span&gt;&lt;span class="line"&gt;&lt;span class="cl"&gt;
&lt;/span&gt;&lt;/span&gt;&lt;span class="line"&gt;&lt;span class="cl"&gt;&lt;span class="c1"&gt;// If endpoint is behind us, move it to the near plane
&lt;/span&gt;&lt;/span&gt;&lt;/span&gt;&lt;span class="line"&gt;&lt;span class="cl"&gt;&lt;span class="k"&gt;if&lt;/span&gt; &lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="nx"&gt;cy1&lt;/span&gt; &lt;span class="o"&gt;&amp;lt;&lt;/span&gt; &lt;span class="nx"&gt;NEAR_PLANE&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt;
&lt;/span&gt;&lt;/span&gt;&lt;span class="line"&gt;&lt;span class="cl"&gt; &lt;span class="kr"&gt;const&lt;/span&gt; &lt;span class="nx"&gt;t&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="nx"&gt;NEAR_PLANE&lt;/span&gt; &lt;span class="o"&gt;-&lt;/span&gt; &lt;span class="nx"&gt;cy1&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt; &lt;span class="o"&gt;/&lt;/span&gt; &lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="nx"&gt;cy2&lt;/span&gt; &lt;span class="o"&gt;-&lt;/span&gt; &lt;span class="nx"&gt;cy1&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt; &lt;span class="c1"&gt;// Interpolation factor
&lt;/span&gt;&lt;/span&gt;&lt;/span&gt;&lt;span class="line"&gt;&lt;span class="cl"&gt; &lt;span class="nx"&gt;cx1&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="nx"&gt;cx1&lt;/span&gt; &lt;span class="o"&gt;+&lt;/span&gt; &lt;span class="nx"&gt;t&lt;/span&gt; &lt;span class="o"&gt;*&lt;/span&gt; &lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="nx"&gt;cx2&lt;/span&gt; &lt;span class="o"&gt;-&lt;/span&gt; &lt;span class="nx"&gt;cx1&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt; &lt;span class="c1"&gt;// New X at near plane
&lt;/span&gt;&lt;/span&gt;&lt;/span&gt;&lt;span class="line"&gt;&lt;span class="cl"&gt; &lt;span class="nx"&gt;cy1&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="nx"&gt;NEAR_PLANE&lt;/span&gt; &lt;span class="c1"&gt;// Clamp to near plane
&lt;/span&gt;&lt;/span&gt;&lt;/span&gt;&lt;span class="line"&gt;&lt;span class="cl"&gt;&lt;span class="p"&gt;}&lt;/span&gt;
&lt;/span&gt;&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;&lt;/div&gt;&lt;h3 id="screen-space-clipping"&gt;Screen-Space Clipping&lt;a class="heading-anchor" href="#screen-space-clipping" aria-label="Link to this section"&gt;&lt;/a&gt;
&lt;/h3&gt;
&lt;p&gt;Walls extending off-screen are also clipped to prevent wasted drawing.&lt;/p&gt;
&lt;div class="highlight"&gt;&lt;pre tabindex="0" class="chroma"&gt;&lt;code class="language-javascript" data-lang="javascript"&gt;&lt;span class="line"&gt;&lt;span class="cl"&gt;&lt;span class="c1"&gt;// Clip to left edge (x = 0)
&lt;/span&gt;&lt;/span&gt;&lt;/span&gt;&lt;span class="line"&gt;&lt;span class="cl"&gt;&lt;span class="k"&gt;if&lt;/span&gt; &lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="nx"&gt;sx1&lt;/span&gt; &lt;span class="o"&gt;&amp;lt;&lt;/span&gt; &lt;span class="mi"&gt;0&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt;
&lt;/span&gt;&lt;/span&gt;&lt;span class="line"&gt;&lt;span class="cl"&gt; &lt;span class="kr"&gt;const&lt;/span&gt; &lt;span class="nx"&gt;t&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="mi"&gt;0&lt;/span&gt; &lt;span class="o"&gt;-&lt;/span&gt; &lt;span class="nx"&gt;sx1&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt; &lt;span class="o"&gt;/&lt;/span&gt; &lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="nx"&gt;sx2&lt;/span&gt; &lt;span class="o"&gt;-&lt;/span&gt; &lt;span class="nx"&gt;sx1&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
&lt;/span&gt;&lt;/span&gt;&lt;span class="line"&gt;&lt;span class="cl"&gt; &lt;span class="nx"&gt;wallTop1&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="nx"&gt;wallTop1&lt;/span&gt; &lt;span class="o"&gt;+&lt;/span&gt; &lt;span class="nx"&gt;t&lt;/span&gt; &lt;span class="o"&gt;*&lt;/span&gt; &lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="nx"&gt;wallTop2&lt;/span&gt; &lt;span class="o"&gt;-&lt;/span&gt; &lt;span class="nx"&gt;wallTop1&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
&lt;/span&gt;&lt;/span&gt;&lt;span class="line"&gt;&lt;span class="cl"&gt; &lt;span class="nx"&gt;wallBot1&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="nx"&gt;wallBot1&lt;/span&gt; &lt;span class="o"&gt;+&lt;/span&gt; &lt;span class="nx"&gt;t&lt;/span&gt; &lt;span class="o"&gt;*&lt;/span&gt; &lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="nx"&gt;wallBot2&lt;/span&gt; &lt;span class="o"&gt;-&lt;/span&gt; &lt;span class="nx"&gt;wallBot1&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
&lt;/span&gt;&lt;/span&gt;&lt;span class="line"&gt;&lt;span class="cl"&gt; &lt;span class="nx"&gt;sx1&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="mi"&gt;0&lt;/span&gt;
&lt;/span&gt;&lt;/span&gt;&lt;span class="line"&gt;&lt;span class="cl"&gt;&lt;span class="p"&gt;}&lt;/span&gt;
&lt;/span&gt;&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;&lt;/div&gt;&lt;h2 id="movement--controls"&gt;Movement &amp;amp; Controls&lt;a class="heading-anchor" href="#movement--controls" aria-label="Link to this section"&gt;&lt;/a&gt;
&lt;/h2&gt;
&lt;h3 id="wasd-movement"&gt;WASD Movement&lt;a class="heading-anchor" href="#wasd-movement" aria-label="Link to this section"&gt;&lt;/a&gt;
&lt;/h3&gt;
&lt;p&gt;Movement is relative to where the player is looking.&lt;/p&gt;
&lt;div class="highlight"&gt;&lt;pre tabindex="0" class="chroma"&gt;&lt;code class="language-javascript" data-lang="javascript"&gt;&lt;span class="line"&gt;&lt;span class="cl"&gt;&lt;span class="kd"&gt;function&lt;/span&gt; &lt;span class="nx"&gt;movePlayer&lt;/span&gt;&lt;span class="p"&gt;()&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt;
&lt;/span&gt;&lt;/span&gt;&lt;span class="line"&gt;&lt;span class="cl"&gt; &lt;span class="c1"&gt;// Forward direction based on angle
&lt;/span&gt;&lt;/span&gt;&lt;/span&gt;&lt;span class="line"&gt;&lt;span class="cl"&gt; &lt;span class="kr"&gt;const&lt;/span&gt; &lt;span class="nx"&gt;dx&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="o"&gt;-&lt;/span&gt;&lt;span class="nb"&gt;Math&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nx"&gt;sin&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="nx"&gt;DEG_TO_RAD&lt;/span&gt;&lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="nx"&gt;Player&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nx"&gt;angle&lt;/span&gt;&lt;span class="p"&gt;])&lt;/span&gt; &lt;span class="o"&gt;*&lt;/span&gt; &lt;span class="nx"&gt;MOVE_SPEED&lt;/span&gt;
&lt;/span&gt;&lt;/span&gt;&lt;span class="line"&gt;&lt;span class="cl"&gt; &lt;span class="kr"&gt;const&lt;/span&gt; &lt;span class="nx"&gt;dy&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="nb"&gt;Math&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nx"&gt;cos&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="nx"&gt;DEG_TO_RAD&lt;/span&gt;&lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="nx"&gt;Player&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nx"&gt;angle&lt;/span&gt;&lt;span class="p"&gt;])&lt;/span&gt; &lt;span class="o"&gt;*&lt;/span&gt; &lt;span class="nx"&gt;MOVE_SPEED&lt;/span&gt;
&lt;/span&gt;&lt;/span&gt;&lt;span class="line"&gt;&lt;span class="cl"&gt; 
&lt;/span&gt;&lt;/span&gt;&lt;span class="line"&gt;&lt;span class="cl"&gt; &lt;span class="c1"&gt;// ... apply movement based on keys ...
&lt;/span&gt;&lt;/span&gt;&lt;/span&gt;&lt;span class="line"&gt;&lt;span class="cl"&gt; 
&lt;/span&gt;&lt;/span&gt;&lt;span class="line"&gt;&lt;span class="cl"&gt; &lt;span class="c1"&gt;// Normalize diagonal movement for consistent speed
&lt;/span&gt;&lt;/span&gt;&lt;/span&gt;&lt;span class="line"&gt;&lt;span class="cl"&gt; &lt;span class="kr"&gt;const&lt;/span&gt; &lt;span class="nx"&gt;magnitude&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="nb"&gt;Math&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nx"&gt;sqrt&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="nx"&gt;moveX&lt;/span&gt;&lt;span class="o"&gt;**&lt;/span&gt;&lt;span class="mi"&gt;2&lt;/span&gt; &lt;span class="o"&gt;+&lt;/span&gt; &lt;span class="nx"&gt;moveY&lt;/span&gt;&lt;span class="o"&gt;**&lt;/span&gt;&lt;span class="mi"&gt;2&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
&lt;/span&gt;&lt;/span&gt;&lt;span class="line"&gt;&lt;span class="cl"&gt; &lt;span class="k"&gt;if&lt;/span&gt; &lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="nx"&gt;magnitude&lt;/span&gt; &lt;span class="o"&gt;&amp;gt;&lt;/span&gt; &lt;span class="mi"&gt;0&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt;
&lt;/span&gt;&lt;/span&gt;&lt;span class="line"&gt;&lt;span class="cl"&gt; &lt;span class="nx"&gt;moveX&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="nx"&gt;moveX&lt;/span&gt; &lt;span class="o"&gt;/&lt;/span&gt; &lt;span class="nx"&gt;magnitude&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt; &lt;span class="o"&gt;*&lt;/span&gt; &lt;span class="nx"&gt;MOVE_SPEED&lt;/span&gt;
&lt;/span&gt;&lt;/span&gt;&lt;span class="line"&gt;&lt;span class="cl"&gt; &lt;span class="nx"&gt;moveY&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="nx"&gt;moveY&lt;/span&gt; &lt;span class="o"&gt;/&lt;/span&gt; &lt;span class="nx"&gt;magnitude&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt; &lt;span class="o"&gt;*&lt;/span&gt; &lt;span class="nx"&gt;MOVE_SPEED&lt;/span&gt;
&lt;/span&gt;&lt;/span&gt;&lt;span class="line"&gt;&lt;span class="cl"&gt; &lt;span class="p"&gt;}&lt;/span&gt;
&lt;/span&gt;&lt;/span&gt;&lt;span class="line"&gt;&lt;span class="cl"&gt;&lt;span class="p"&gt;}&lt;/span&gt;
&lt;/span&gt;&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;&lt;/div&gt;&lt;h3 id="mouse-look-with-pointer-lock"&gt;Mouse Look with Pointer Lock&lt;a class="heading-anchor" href="#mouse-look-with-pointer-lock" aria-label="Link to this section"&gt;&lt;/a&gt;
&lt;/h3&gt;
&lt;p&gt;For immersive FPS-style controls, we lock the mouse cursor. &lt;strong&gt;Pointer lock&lt;/strong&gt; hides the cursor and provides raw mouse movement data.&lt;/p&gt;
&lt;div class="highlight"&gt;&lt;pre tabindex="0" class="chroma"&gt;&lt;code class="language-javascript" data-lang="javascript"&gt;&lt;span class="line"&gt;&lt;span class="cl"&gt;&lt;span class="nb"&gt;document&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nx"&gt;addEventListener&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="s2"&gt;&amp;#34;mousemove&amp;#34;&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="nx"&gt;event&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt; &lt;span class="p"&gt;=&amp;gt;&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt;
&lt;/span&gt;&lt;/span&gt;&lt;span class="line"&gt;&lt;span class="cl"&gt; &lt;span class="k"&gt;if&lt;/span&gt; &lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="o"&gt;!&lt;/span&gt;&lt;span class="nx"&gt;isMouseLocked&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt; &lt;span class="k"&gt;return&lt;/span&gt;
&lt;/span&gt;&lt;/span&gt;&lt;span class="line"&gt;&lt;span class="cl"&gt; 
&lt;/span&gt;&lt;/span&gt;&lt;span class="line"&gt;&lt;span class="cl"&gt; &lt;span class="kr"&gt;const&lt;/span&gt; &lt;span class="nx"&gt;mouseMove&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="nx"&gt;event&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nx"&gt;movementX&lt;/span&gt;
&lt;/span&gt;&lt;/span&gt;&lt;span class="line"&gt;&lt;span class="cl"&gt; &lt;span class="kr"&gt;const&lt;/span&gt; &lt;span class="nx"&gt;degreeChange&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="nx"&gt;mouseMove&lt;/span&gt; &lt;span class="o"&gt;/&lt;/span&gt; &lt;span class="nx"&gt;screen&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nx"&gt;width&lt;/span&gt; &lt;span class="o"&gt;*&lt;/span&gt; &lt;span class="mi"&gt;360&lt;/span&gt;
&lt;/span&gt;&lt;/span&gt;&lt;span class="line"&gt;&lt;span class="cl"&gt; 
&lt;/span&gt;&lt;/span&gt;&lt;span class="line"&gt;&lt;span class="cl"&gt; &lt;span class="nx"&gt;Player&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nx"&gt;angle&lt;/span&gt; &lt;span class="o"&gt;-=&lt;/span&gt; &lt;span class="nx"&gt;degreeChange&lt;/span&gt;
&lt;/span&gt;&lt;/span&gt;&lt;span class="line"&gt;&lt;span class="cl"&gt;&lt;span class="p"&gt;})&lt;/span&gt;
&lt;/span&gt;&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;&lt;/div&gt;&lt;h2 id="wall-collision-detection"&gt;Wall Collision Detection&lt;a class="heading-anchor" href="#wall-collision-detection" aria-label="Link to this section"&gt;&lt;/a&gt;
&lt;/h2&gt;
&lt;h3 id="grid-based-collision"&gt;Grid-Based Collision&lt;a class="heading-anchor" href="#grid-based-collision" aria-label="Link to this section"&gt;&lt;/a&gt;
&lt;/h3&gt;
&lt;p&gt;Since our world is grid-based, collision is simple. Just check if the players future calculated position is a wall and conditionally allow movement.&lt;/p&gt;
&lt;div class="highlight"&gt;&lt;pre tabindex="0" class="chroma"&gt;&lt;code class="language-javascript" data-lang="javascript"&gt;&lt;span class="line"&gt;&lt;span class="cl"&gt;&lt;span class="kd"&gt;function&lt;/span&gt; &lt;span class="nx"&gt;isWall&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="nx"&gt;worldX&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="nx"&gt;worldY&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt;
&lt;/span&gt;&lt;/span&gt;&lt;span class="line"&gt;&lt;span class="cl"&gt; &lt;span class="kr"&gt;const&lt;/span&gt; &lt;span class="nx"&gt;tileX&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="nb"&gt;Math&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nx"&gt;floor&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="nx"&gt;worldX&lt;/span&gt; &lt;span class="o"&gt;/&lt;/span&gt; &lt;span class="nx"&gt;TILE_SIZE&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
&lt;/span&gt;&lt;/span&gt;&lt;span class="line"&gt;&lt;span class="cl"&gt; &lt;span class="kr"&gt;const&lt;/span&gt; &lt;span class="nx"&gt;tileY&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="nb"&gt;Math&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nx"&gt;floor&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="nx"&gt;worldY&lt;/span&gt; &lt;span class="o"&gt;/&lt;/span&gt; &lt;span class="nx"&gt;TILE_SIZE&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
&lt;/span&gt;&lt;/span&gt;&lt;span class="line"&gt;&lt;span class="cl"&gt; &lt;span class="k"&gt;return&lt;/span&gt; &lt;span class="nx"&gt;getMapTile&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="nx"&gt;tileX&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="nx"&gt;tileY&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt; &lt;span class="o"&gt;===&lt;/span&gt; &lt;span class="mi"&gt;1&lt;/span&gt; &lt;span class="c1"&gt;// 1 is a wall in our implementation
&lt;/span&gt;&lt;/span&gt;&lt;/span&gt;&lt;span class="line"&gt;&lt;span class="cl"&gt;&lt;span class="p"&gt;}&lt;/span&gt;
&lt;/span&gt;&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;&lt;/div&gt;&lt;p&gt;We also give the player a &lt;strong&gt;radius&lt;/strong&gt; and check all corners to ensure the player doesn&amp;rsquo;t clip into walls. If directly blocked, we try sliding along the wall. This creates smooth movement even when running into walls at angles.&lt;/p&gt;
&lt;h2 id="putting-it-all-together"&gt;Putting It All Together&lt;a class="heading-anchor" href="#putting-it-all-together" aria-label="Link to this section"&gt;&lt;/a&gt;
&lt;/h2&gt;
&lt;p&gt;Every frame, we simply clear the screen, handle input, render the walls, and debug if needed.&lt;/p&gt;
&lt;div class="highlight"&gt;&lt;pre tabindex="0" class="chroma"&gt;&lt;code class="language-javascript" data-lang="javascript"&gt;&lt;span class="line"&gt;&lt;span class="cl"&gt;&lt;span class="kd"&gt;function&lt;/span&gt; &lt;span class="nx"&gt;loop&lt;/span&gt;&lt;span class="p"&gt;()&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt;
&lt;/span&gt;&lt;/span&gt;&lt;span class="line"&gt;&lt;span class="cl"&gt; &lt;span class="nx"&gt;clearScreen&lt;/span&gt;&lt;span class="p"&gt;()&lt;/span&gt; &lt;span class="c1"&gt;// Draw floor and ceiling
&lt;/span&gt;&lt;/span&gt;&lt;/span&gt;&lt;span class="line"&gt;&lt;span class="cl"&gt; &lt;span class="nx"&gt;movePlayer&lt;/span&gt;&lt;span class="p"&gt;()&lt;/span&gt; &lt;span class="c1"&gt;// Handle input
&lt;/span&gt;&lt;/span&gt;&lt;/span&gt;&lt;span class="line"&gt;&lt;span class="cl"&gt; &lt;span class="nx"&gt;draw3D&lt;/span&gt;&lt;span class="p"&gt;()&lt;/span&gt; &lt;span class="c1"&gt;// Render walls
&lt;/span&gt;&lt;/span&gt;&lt;/span&gt;&lt;span class="line"&gt;&lt;span class="cl"&gt; &lt;span class="nx"&gt;debug&lt;/span&gt;&lt;span class="p"&gt;()&lt;/span&gt; &lt;span class="c1"&gt;// Show debug info
&lt;/span&gt;&lt;/span&gt;&lt;/span&gt;&lt;span class="line"&gt;&lt;span class="cl"&gt;&lt;span class="p"&gt;}&lt;/span&gt;
&lt;/span&gt;&lt;/span&gt;&lt;span class="line"&gt;&lt;span class="cl"&gt;
&lt;/span&gt;&lt;/span&gt;&lt;span class="line"&gt;&lt;span class="cl"&gt;&lt;span class="c1"&gt;// Run at 30 FPS
&lt;/span&gt;&lt;/span&gt;&lt;/span&gt;&lt;span class="line"&gt;&lt;span class="cl"&gt;&lt;span class="nx"&gt;setInterval&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="nx"&gt;loop&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="mi"&gt;1000&lt;/span&gt; &lt;span class="o"&gt;/&lt;/span&gt; &lt;span class="mi"&gt;30&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
&lt;/span&gt;&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;&lt;/div&gt;&lt;h2 id="what-next"&gt;What Next?&lt;a class="heading-anchor" href="#what-next" aria-label="Link to this section"&gt;&lt;/a&gt;
&lt;/h2&gt;
&lt;p&gt;The engine we built covers the fundamentals, but classic DOOM had much more: texture mapping, sprites, variable height sectors for different floor/ceiling heights, lighting, enemy AI, sound, guns, an actual game!&lt;/p&gt;
&lt;p&gt;Feel free to extend this engine to add more features! &amp;lt;3&lt;/p&gt;
&lt;h2 id="references"&gt;References&lt;a class="heading-anchor" href="#references" aria-label="Link to this section"&gt;&lt;/a&gt;
&lt;/h2&gt;
&lt;ul&gt;
&lt;li&gt;&lt;a href="https://doomwiki.org/wiki/Rendering_engine" data-external&gt;DOOM Wiki: Rendering Engine&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;a href="https://www.youtube.com/watch?v=huMO4VQEwPc" data-external&gt;Let&amp;rsquo;s program DOOM by 3DSage&lt;/a&gt;&lt;/li&gt;
&lt;/ul&gt;</content:encoded></item><item><title>Publish to NPM from GitHub Actions using OIDC</title><link>https://ankush.one/blogs/npm-oidc-publishing/</link><pubDate>Tue, 09 Dec 2025 00:00:00 +0000</pubDate><author>ankush4singh@gmail.com (Ankush Singh)</author><guid>https://ankush.one/blogs/npm-oidc-publishing/</guid><category>developer-tools</category><category>tutorial</category><description>Learn how to automate npm package publishing through GitHub Actions using OpenID Connect (OIDC) - no API keys or access tokens required.</description><content:encoded>&lt;p&gt;I frequently create developer tools and publish node packages on npmjs, but recently classic tokens got deprecated, granular tokens are limited to 90 days, require 2FA and a bunch of changes were introduced to improve security around publishing packages. More about that on &lt;a href="https://github.blog/changelog/2025-12-09-npm-classic-tokens-revoked-session-based-auth-and-cli-token-management-now-available/" data-external&gt;this GitHub announcement&lt;/a&gt;.&lt;/p&gt;
&lt;p&gt;One of the improvements was the introduction of &lt;strong&gt;&lt;a href="https://openid.net/developers/how-connect-works/" data-external&gt;OpenID Connect&lt;/a&gt; (OIDC)&lt;/strong&gt;, which allows developers to publish npm packages through GitHub Actions, without having to setup any kind of API key or access token.&lt;/p&gt;
&lt;p&gt;Today we&amp;rsquo;re having a look at how you can implement it for your packages as well.&lt;/p&gt;
&lt;h2 id="step-1-setup-a-node-package-or-skip-if-you-already-have-one"&gt;Step 1: Setup a Node Package (or skip if you already have one)&lt;a class="heading-anchor" href="#step-1-setup-a-node-package-or-skip-if-you-already-have-one" aria-label="Link to this section"&gt;&lt;/a&gt;
&lt;/h2&gt;
&lt;p&gt;As an example, we&amp;rsquo;ll create a fresh package which just logs &amp;ldquo;hello world&amp;rdquo; to console.&lt;/p&gt;
&lt;aside class="doc-callout doc-callout--note" aria-label="Note"&gt;
&lt;p class="doc-callout-title"&gt;&lt;span class="doc-callout-icon"&gt;&lt;svg class="ui-glyph ui-glyph-note" width="1em" height="1em" viewBox="0 0 16 16" fill="none" stroke="currentColor" stroke-width="1.5" stroke-linecap="round" stroke-linejoin="round" aria-hidden="true" focusable="false"&gt;&lt;circle cx="8" cy="8" r="6"/&gt;&lt;path d="M8 7.25v4M8 4.75v.01"/&gt;&lt;/svg&gt;&lt;/span&gt;&lt;span&gt;Note&lt;/span&gt;&lt;/p&gt;
&lt;div class="doc-callout-body"&gt;
&lt;p&gt;This package is only meant as a demo for this blog. It is not recommended to create such packages for yourself — you should just setup OIDC on your existing packages :)&lt;/p&gt;
&lt;/div&gt;
&lt;/aside&gt;&lt;div class="highlight"&gt;&lt;pre tabindex="0" class="chroma"&gt;&lt;code class="language-bash" data-lang="bash"&gt;&lt;span class="line"&gt;&lt;span class="cl"&gt;mkdir npm-oidc-demo &lt;span class="o"&gt;&amp;amp;&amp;amp;&lt;/span&gt; &lt;span class="nb"&gt;cd&lt;/span&gt; npm-oidc-demo
&lt;/span&gt;&lt;/span&gt;&lt;span class="line"&gt;&lt;span class="cl"&gt;npm init -y
&lt;/span&gt;&lt;/span&gt;&lt;span class="line"&gt;&lt;span class="cl"&gt;touch index.js
&lt;/span&gt;&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;&lt;/div&gt;&lt;p&gt;&lt;strong&gt;index.js:&lt;/strong&gt;&lt;/p&gt;
&lt;div class="highlight"&gt;&lt;pre tabindex="0" class="chroma"&gt;&lt;code class="language-javascript" data-lang="javascript"&gt;&lt;span class="line"&gt;&lt;span class="cl"&gt;&lt;span class="kr"&gt;export&lt;/span&gt; &lt;span class="kd"&gt;function&lt;/span&gt; &lt;span class="nx"&gt;helloWorld&lt;/span&gt;&lt;span class="p"&gt;()&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt;
&lt;/span&gt;&lt;/span&gt;&lt;span class="line"&gt;&lt;span class="cl"&gt; &lt;span class="nx"&gt;console&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nx"&gt;log&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="s2"&gt;&amp;#34;Hello World&amp;#34;&lt;/span&gt;&lt;span class="p"&gt;);&lt;/span&gt;
&lt;/span&gt;&lt;/span&gt;&lt;span class="line"&gt;&lt;span class="cl"&gt;&lt;span class="p"&gt;}&lt;/span&gt;
&lt;/span&gt;&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;&lt;/div&gt;&lt;p&gt;&lt;strong&gt;package.json:&lt;/strong&gt;&lt;/p&gt;
&lt;div class="highlight"&gt;&lt;pre tabindex="0" class="chroma"&gt;&lt;code class="language-json" data-lang="json"&gt;&lt;span class="line"&gt;&lt;span class="cl"&gt;&lt;span class="p"&gt;{&lt;/span&gt;
&lt;/span&gt;&lt;/span&gt;&lt;span class="line"&gt;&lt;span class="cl"&gt; &lt;span class="nt"&gt;&amp;#34;name&amp;#34;&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="s2"&gt;&amp;#34;npm-oidc-demo&amp;#34;&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
&lt;/span&gt;&lt;/span&gt;&lt;span class="line"&gt;&lt;span class="cl"&gt; &lt;span class="nt"&gt;&amp;#34;version&amp;#34;&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="s2"&gt;&amp;#34;1.0.1&amp;#34;&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
&lt;/span&gt;&lt;/span&gt;&lt;span class="line"&gt;&lt;span class="cl"&gt; &lt;span class="nt"&gt;&amp;#34;description&amp;#34;&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="s2"&gt;&amp;#34;A simple hello world package&amp;#34;&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
&lt;/span&gt;&lt;/span&gt;&lt;span class="line"&gt;&lt;span class="cl"&gt; &lt;span class="nt"&gt;&amp;#34;main&amp;#34;&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="s2"&gt;&amp;#34;index.js&amp;#34;&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
&lt;/span&gt;&lt;/span&gt;&lt;span class="line"&gt;&lt;span class="cl"&gt; &lt;span class="nt"&gt;&amp;#34;type&amp;#34;&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="s2"&gt;&amp;#34;module&amp;#34;&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
&lt;/span&gt;&lt;/span&gt;&lt;span class="line"&gt;&lt;span class="cl"&gt; &lt;span class="nt"&gt;&amp;#34;exports&amp;#34;&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt;
&lt;/span&gt;&lt;/span&gt;&lt;span class="line"&gt;&lt;span class="cl"&gt; &lt;span class="nt"&gt;&amp;#34;.&amp;#34;&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="s2"&gt;&amp;#34;./index.js&amp;#34;&lt;/span&gt;
&lt;/span&gt;&lt;/span&gt;&lt;span class="line"&gt;&lt;span class="cl"&gt; &lt;span class="p"&gt;},&lt;/span&gt;
&lt;/span&gt;&lt;/span&gt;&lt;span class="line"&gt;&lt;span class="cl"&gt; &lt;span class="nt"&gt;&amp;#34;repository&amp;#34;&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt;
&lt;/span&gt;&lt;/span&gt;&lt;span class="line"&gt;&lt;span class="cl"&gt; &lt;span class="nt"&gt;&amp;#34;type&amp;#34;&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="s2"&gt;&amp;#34;git&amp;#34;&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
&lt;/span&gt;&lt;/span&gt;&lt;span class="line"&gt;&lt;span class="cl"&gt; &lt;span class="nt"&gt;&amp;#34;url&amp;#34;&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="s2"&gt;&amp;#34;https://github.com/ankushKun/npm-oidc-demo&amp;#34;&lt;/span&gt;
&lt;/span&gt;&lt;/span&gt;&lt;span class="line"&gt;&lt;span class="cl"&gt; &lt;span class="p"&gt;},&lt;/span&gt;
&lt;/span&gt;&lt;/span&gt;&lt;span class="line"&gt;&lt;span class="cl"&gt; &lt;span class="nt"&gt;&amp;#34;keywords&amp;#34;&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="p"&gt;[&lt;/span&gt;
&lt;/span&gt;&lt;/span&gt;&lt;span class="line"&gt;&lt;span class="cl"&gt; &lt;span class="s2"&gt;&amp;#34;npm&amp;#34;&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
&lt;/span&gt;&lt;/span&gt;&lt;span class="line"&gt;&lt;span class="cl"&gt; &lt;span class="s2"&gt;&amp;#34;oidc&amp;#34;&lt;/span&gt;
&lt;/span&gt;&lt;/span&gt;&lt;span class="line"&gt;&lt;span class="cl"&gt; &lt;span class="p"&gt;],&lt;/span&gt;
&lt;/span&gt;&lt;/span&gt;&lt;span class="line"&gt;&lt;span class="cl"&gt; &lt;span class="nt"&gt;&amp;#34;author&amp;#34;&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="s2"&gt;&amp;#34;&amp;#34;&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
&lt;/span&gt;&lt;/span&gt;&lt;span class="line"&gt;&lt;span class="cl"&gt; &lt;span class="nt"&gt;&amp;#34;license&amp;#34;&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="s2"&gt;&amp;#34;ISC&amp;#34;&lt;/span&gt;
&lt;/span&gt;&lt;/span&gt;&lt;span class="line"&gt;&lt;span class="cl"&gt;&lt;span class="p"&gt;}&lt;/span&gt;
&lt;/span&gt;&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;&lt;/div&gt;&lt;aside class="doc-callout doc-callout--note" aria-label="Note"&gt;
&lt;p class="doc-callout-title"&gt;&lt;span class="doc-callout-icon"&gt;&lt;svg class="ui-glyph ui-glyph-note" width="1em" height="1em" viewBox="0 0 16 16" fill="none" stroke="currentColor" stroke-width="1.5" stroke-linecap="round" stroke-linejoin="round" aria-hidden="true" focusable="false"&gt;&lt;circle cx="8" cy="8" r="6"/&gt;&lt;path d="M8 7.25v4M8 4.75v.01"/&gt;&lt;/svg&gt;&lt;/span&gt;&lt;span&gt;Note&lt;/span&gt;&lt;/p&gt;
&lt;div class="doc-callout-body"&gt;
&lt;p&gt;An already published package is required to use OIDC (might change in the future, who knows…)&lt;/p&gt;
&lt;/div&gt;
&lt;/aside&gt;&lt;p&gt;Run &lt;code&gt;npm login&lt;/code&gt; and follow the steps to link the npm CLI with your npmjs account. When successful, run &lt;code&gt;npm whoami&lt;/code&gt; and it should print your username.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Ready for publishing:&lt;/strong&gt;&lt;/p&gt;
&lt;blockquote&gt;
&lt;p&gt;If you ignored the previous note and still followed this guide to publish the demo package yourself, you might need to change your package name as I have already used “npm-oidc-demo”&lt;/p&gt;

&lt;/blockquote&gt;&lt;div class="highlight"&gt;&lt;pre tabindex="0" class="chroma"&gt;&lt;code class="language-bash" data-lang="bash"&gt;&lt;span class="line"&gt;&lt;span class="cl"&gt;$ npm whoami
&lt;/span&gt;&lt;/span&gt;&lt;span class="line"&gt;&lt;span class="cl"&gt;ankushkun
&lt;/span&gt;&lt;/span&gt;&lt;span class="line"&gt;&lt;span class="cl"&gt;
&lt;/span&gt;&lt;/span&gt;&lt;span class="line"&gt;&lt;span class="cl"&gt;$ npm publish
&lt;/span&gt;&lt;/span&gt;&lt;span class="line"&gt;&lt;span class="cl"&gt;npm notice 
&lt;/span&gt;&lt;/span&gt;&lt;span class="line"&gt;&lt;span class="cl"&gt;npm notice 📦 npm-oidc-demo@1.0.0
&lt;/span&gt;&lt;/span&gt;&lt;span class="line"&gt;&lt;span class="cl"&gt;...
&lt;/span&gt;&lt;/span&gt;&lt;span class="line"&gt;&lt;span class="cl"&gt;+ npm-oidc-demo@1.0.0
&lt;/span&gt;&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;&lt;/div&gt;&lt;p&gt;The package is now available at &lt;a href="https://www.npmjs.com/package/npm-oidc-demo" data-external&gt;&lt;code&gt;https://www.npmjs.com/package/npm-oidc-demo&lt;/code&gt;&lt;/a&gt;.&lt;/p&gt;
&lt;h2 id="step-2-update-oidc-details-on-package-settings"&gt;Step 2: Update OIDC Details on Package Settings&lt;a class="heading-anchor" href="#step-2-update-oidc-details-on-package-settings" aria-label="Link to this section"&gt;&lt;/a&gt;
&lt;/h2&gt;
&lt;p&gt;Head over to the settings tab on your freshly published package and select &lt;strong&gt;GitHub Actions&lt;/strong&gt; as the trusted publisher.&lt;/p&gt;
&lt;figure class="doc-figure"&gt;&lt;img src="https://miro.medium.com/v2/resize:fit:1400/format:webp/1*s0fSghe8DIxQPw6K-OhozA.png" alt="npmjs package settings tab" loading="lazy" decoding="async"&gt;&lt;figcaption aria-hidden="true"&gt;npmjs package settings tab&lt;/figcaption&gt;&lt;/figure&gt;&lt;figure class="doc-figure"&gt;&lt;img src="https://miro.medium.com/v2/resize:fit:1400/format:webp/1*4Fc2oJ3_Ijr8uLga0P4v_A.png" alt="GitHub actions trusted publisher settings" loading="lazy" decoding="async"&gt;&lt;figcaption aria-hidden="true"&gt;GitHub actions trusted publisher settings&lt;/figcaption&gt;&lt;/figure&gt;&lt;p&gt;I have already created a GitHub repo at &lt;a href="https://github.com/ankushKun/npm-oidc-demo" data-external&gt;&lt;code&gt;ankushKun/npm-oidc-demo&lt;/code&gt;&lt;/a&gt;, now we just need to create a workflow file. Put this inside &lt;code&gt;.github/workflows/publish.yml&lt;/code&gt;:&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;publish.yml:&lt;/strong&gt;&lt;/p&gt;
&lt;div class="highlight"&gt;&lt;pre tabindex="0" class="chroma"&gt;&lt;code class="language-yaml" data-lang="yaml"&gt;&lt;span class="line"&gt;&lt;span class="cl"&gt;&lt;span class="nt"&gt;name&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="l"&gt;Publish to npm&lt;/span&gt;&lt;span class="w"&gt;
&lt;/span&gt;&lt;/span&gt;&lt;/span&gt;&lt;span class="line"&gt;&lt;span class="cl"&gt;&lt;span class="w"&gt;
&lt;/span&gt;&lt;/span&gt;&lt;/span&gt;&lt;span class="line"&gt;&lt;span class="cl"&gt;&lt;span class="nt"&gt;on&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt;
&lt;/span&gt;&lt;/span&gt;&lt;/span&gt;&lt;span class="line"&gt;&lt;span class="cl"&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="nt"&gt;release&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt;
&lt;/span&gt;&lt;/span&gt;&lt;/span&gt;&lt;span class="line"&gt;&lt;span class="cl"&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="nt"&gt;types&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="l"&gt;published]&lt;/span&gt;&lt;span class="w"&gt;
&lt;/span&gt;&lt;/span&gt;&lt;/span&gt;&lt;span class="line"&gt;&lt;span class="cl"&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="nt"&gt;workflow_dispatch&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt;
&lt;/span&gt;&lt;/span&gt;&lt;/span&gt;&lt;span class="line"&gt;&lt;span class="cl"&gt;&lt;span class="w"&gt;
&lt;/span&gt;&lt;/span&gt;&lt;/span&gt;&lt;span class="line"&gt;&lt;span class="cl"&gt;&lt;span class="nt"&gt;jobs&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt;
&lt;/span&gt;&lt;/span&gt;&lt;/span&gt;&lt;span class="line"&gt;&lt;span class="cl"&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="nt"&gt;publish&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt;
&lt;/span&gt;&lt;/span&gt;&lt;/span&gt;&lt;span class="line"&gt;&lt;span class="cl"&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="nt"&gt;runs-on&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="l"&gt;ubuntu-latest&lt;/span&gt;&lt;span class="w"&gt;
&lt;/span&gt;&lt;/span&gt;&lt;/span&gt;&lt;span class="line"&gt;&lt;span class="cl"&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="nt"&gt;permissions&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt;
&lt;/span&gt;&lt;/span&gt;&lt;/span&gt;&lt;span class="line"&gt;&lt;span class="cl"&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="nt"&gt;contents&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="l"&gt;read&lt;/span&gt;&lt;span class="w"&gt;
&lt;/span&gt;&lt;/span&gt;&lt;/span&gt;&lt;span class="line"&gt;&lt;span class="cl"&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="nt"&gt;id-token&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="l"&gt;write&lt;/span&gt;&lt;span class="w"&gt;
&lt;/span&gt;&lt;/span&gt;&lt;/span&gt;&lt;span class="line"&gt;&lt;span class="cl"&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="nt"&gt;steps&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt;
&lt;/span&gt;&lt;/span&gt;&lt;/span&gt;&lt;span class="line"&gt;&lt;span class="cl"&gt;&lt;span class="w"&gt; &lt;/span&gt;- &lt;span class="nt"&gt;uses&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="l"&gt;actions/checkout@v4&lt;/span&gt;&lt;span class="w"&gt;
&lt;/span&gt;&lt;/span&gt;&lt;/span&gt;&lt;span class="line"&gt;&lt;span class="cl"&gt;&lt;span class="w"&gt;
&lt;/span&gt;&lt;/span&gt;&lt;/span&gt;&lt;span class="line"&gt;&lt;span class="cl"&gt;&lt;span class="w"&gt; &lt;/span&gt;- &lt;span class="nt"&gt;uses&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="l"&gt;actions/setup-node@v4&lt;/span&gt;&lt;span class="w"&gt;
&lt;/span&gt;&lt;/span&gt;&lt;/span&gt;&lt;span class="line"&gt;&lt;span class="cl"&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="nt"&gt;with&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt;
&lt;/span&gt;&lt;/span&gt;&lt;/span&gt;&lt;span class="line"&gt;&lt;span class="cl"&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="nt"&gt;node-version&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="s2"&gt;&amp;#34;20.x&amp;#34;&lt;/span&gt;&lt;span class="w"&gt;
&lt;/span&gt;&lt;/span&gt;&lt;/span&gt;&lt;span class="line"&gt;&lt;span class="cl"&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="nt"&gt;registry-url&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="s2"&gt;&amp;#34;https://registry.npmjs.org&amp;#34;&lt;/span&gt;&lt;span class="w"&gt;
&lt;/span&gt;&lt;/span&gt;&lt;/span&gt;&lt;span class="line"&gt;&lt;span class="cl"&gt;&lt;span class="w"&gt;
&lt;/span&gt;&lt;/span&gt;&lt;/span&gt;&lt;span class="line"&gt;&lt;span class="cl"&gt;&lt;span class="w"&gt; &lt;/span&gt;- &lt;span class="nt"&gt;name&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="l"&gt;Update npm&lt;/span&gt;&lt;span class="w"&gt;
&lt;/span&gt;&lt;/span&gt;&lt;/span&gt;&lt;span class="line"&gt;&lt;span class="cl"&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="nt"&gt;run&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="l"&gt;npm install -g npm@latest&lt;/span&gt;&lt;span class="w"&gt;
&lt;/span&gt;&lt;/span&gt;&lt;/span&gt;&lt;span class="line"&gt;&lt;span class="cl"&gt;&lt;span class="w"&gt;
&lt;/span&gt;&lt;/span&gt;&lt;/span&gt;&lt;span class="line"&gt;&lt;span class="cl"&gt;&lt;span class="w"&gt; &lt;/span&gt;- &lt;span class="nt"&gt;name&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="l"&gt;Publish to npm&lt;/span&gt;&lt;span class="w"&gt;
&lt;/span&gt;&lt;/span&gt;&lt;/span&gt;&lt;span class="line"&gt;&lt;span class="cl"&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="nt"&gt;run&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="l"&gt;npm publish&lt;/span&gt;&lt;span class="w"&gt;
&lt;/span&gt;&lt;/span&gt;&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;&lt;/div&gt;&lt;p&gt;Push the YAML file to the repo and update the OIDC details on the package settings. I have kept the environment name empty because I don&amp;rsquo;t have any explicit environment setup for actions.&lt;/p&gt;
&lt;figure class="doc-figure"&gt;&lt;img src="https://miro.medium.com/v2/resize:fit:1400/format:webp/1*fjfYVS4FfV6Av8xqeaMNqw.png" alt="updated oidc details" loading="lazy" decoding="async"&gt;&lt;figcaption aria-hidden="true"&gt;updated oidc details&lt;/figcaption&gt;&lt;/figure&gt;&lt;p&gt;update and save the settings.&lt;/p&gt;
&lt;p&gt;Now when you create a new release on the github repo, a workflow will start running which will publish the package to npm with provenance.&lt;/p&gt;
&lt;h2 id="step-3-create-a-release-on-github-to-run-the-action"&gt;Step 3: Create a Release on GitHub to Run the Action&lt;a class="heading-anchor" href="#step-3-create-a-release-on-github-to-run-the-action" aria-label="Link to this section"&gt;&lt;/a&gt;
&lt;/h2&gt;
&lt;p&gt;Now it&amp;rsquo;s time to trigger the workflow! Go back to the github repo and create a new release.&lt;/p&gt;
&lt;p&gt;Once created, the GitHub Action will automatically start running. You can monitor the progress in the &lt;strong&gt;Actions&lt;/strong&gt; tab of your repository.&lt;/p&gt;
&lt;figure class="doc-figure"&gt;&lt;img src="https://miro.medium.com/v2/resize:fit:1400/format:webp/1*5vdxP7PcaEI51ntX42K5_A.png" alt="successful workflow" loading="lazy" decoding="async"&gt;&lt;figcaption aria-hidden="true"&gt;successful workflow&lt;/figcaption&gt;&lt;/figure&gt;&lt;p&gt;When the workflow completes successfully, your package will be published to npm with provenance — proving it was built and published through your trusted CI/CD pipeline!&lt;/p&gt;
&lt;figure class="doc-figure"&gt;&lt;img src="https://miro.medium.com/v2/resize:fit:1400/format:webp/1*CYJjIDfnPq3HfLTo8G37vA.png" alt="published package" loading="lazy" decoding="async"&gt;&lt;figcaption aria-hidden="true"&gt;published package&lt;/figcaption&gt;&lt;/figure&gt;&lt;p&gt;And that’s how one can setup cicd to publish npm packages from github actions without having to create any access tokens themselves!&lt;/p&gt;</content:encoded></item><item><title>A Noob's Guide to Web APIs</title><link>https://ankush.one/blogs/web-apis-guide/</link><pubDate>Wed, 19 Feb 2025 00:00:00 +0000</pubDate><author>ankush4singh@gmail.com (Ankush Singh)</author><guid>https://ankush.one/blogs/web-apis-guide/</guid><category>web</category><category>tutorial</category><description>Learn what APIs are, how they work, and how to create your own simple API using Node.js and Express.</description><content:encoded>&lt;p&gt;&lt;em&gt;Alllll riiiight lets do this one more time&lt;/em&gt;, my name is Ankush and I&amp;rsquo;m your friendly neighbourhood web developer, here to educate your pathetic ass a bit about web APIs, how they can be used to connect your frontend with your backend and how you can write an API using NodeJS 🫡&lt;/p&gt;
&lt;p&gt;An API — Application Programming Interface, defines how different applications can talk to each other and exchange data through structured requests and responses.&lt;/p&gt;
&lt;h2 id="api-types"&gt;API Types&lt;a class="heading-anchor" href="#api-types" aria-label="Link to this section"&gt;&lt;/a&gt;
&lt;/h2&gt;
&lt;ol&gt;
&lt;li&gt;&lt;strong&gt;Web APIs&lt;/strong&gt; — the one we will be focusing on today&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;OS APIs&lt;/strong&gt; — Windows APIs, Android APIs, etc&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Library APIs&lt;/strong&gt; — Like a python library&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Hardware APIs&lt;/strong&gt; — allow accessing hardware functions (embedded libraries?)&lt;/li&gt;
&lt;/ol&gt;
&lt;h2 id="restful-web-apis"&gt;RESTful Web APIs&lt;a class="heading-anchor" href="#restful-web-apis" aria-label="Link to this section"&gt;&lt;/a&gt;
&lt;/h2&gt;
&lt;p&gt;A RESTful Web API is one that follows the &amp;ldquo;REST&amp;rdquo; (Representational State Transfer) principles, a software architecture style that ensures scalability, simplicity, and stateless communication between servers and clients.&lt;/p&gt;
&lt;h3 id="common-rest-http-methods"&gt;Common REST HTTP Methods&lt;a class="heading-anchor" href="#common-rest-http-methods" aria-label="Link to this section"&gt;&lt;/a&gt;
&lt;/h3&gt;
&lt;p&gt;There are 4 common methods/requests to call REST APIs:&lt;/p&gt;
&lt;ol&gt;
&lt;li&gt;&lt;strong&gt;GET&lt;/strong&gt; — to get some data from the server&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;POST&lt;/strong&gt; — to send/create some data from the server and get a response&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;PUT&lt;/strong&gt; — to update existing data on a server&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;DELETE&lt;/strong&gt; — to remove some data from a server&lt;/li&gt;
&lt;/ol&gt;
&lt;p&gt;When browsing the web, a GET request is used to fetch the client-side content that gets rendered by the browser.&lt;/p&gt;
&lt;h3 id="api-response-codes"&gt;API Response Codes&lt;a class="heading-anchor" href="#api-response-codes" aria-label="Link to this section"&gt;&lt;/a&gt;
&lt;/h3&gt;
&lt;p&gt;Whenever an API is called, it will always return a &lt;strong&gt;status code&lt;/strong&gt; (a number) to signify the result of the request.&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;&lt;strong&gt;1XX&lt;/strong&gt;: Informational response&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;2XX&lt;/strong&gt;: Success codes (e.g., 200 — OK)&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;3XX&lt;/strong&gt;: Redirects (e.g., 301 — Moved Permanently)&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;4XX&lt;/strong&gt;: Client errors (e.g., 400 — Bad Request, 401 — Unauthorised, 404 — Not Found)&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;5XX&lt;/strong&gt;: Server errors (e.g., 500 — Internal Server Error)&lt;/li&gt;
&lt;/ul&gt;
&lt;h2 id="lets-write-an-api-endpoint-ourselves"&gt;Let&amp;rsquo;s Write an API Endpoint Ourselves&lt;a class="heading-anchor" href="#lets-write-an-api-endpoint-ourselves" aria-label="Link to this section"&gt;&lt;/a&gt;
&lt;/h2&gt;
&lt;p&gt;Make sure you have NodeJS + NPM already installed and follow these steps:&lt;/p&gt;
&lt;div class="highlight"&gt;&lt;pre tabindex="0" class="chroma"&gt;&lt;code class="language-bash" data-lang="bash"&gt;&lt;span class="line"&gt;&lt;span class="cl"&gt;mkdir api-test
&lt;/span&gt;&lt;/span&gt;&lt;span class="line"&gt;&lt;span class="cl"&gt;&lt;span class="nb"&gt;cd&lt;/span&gt; api-test
&lt;/span&gt;&lt;/span&gt;&lt;span class="line"&gt;&lt;span class="cl"&gt;npm init -y
&lt;/span&gt;&lt;/span&gt;&lt;span class="line"&gt;&lt;span class="cl"&gt;npm i express
&lt;/span&gt;&lt;/span&gt;&lt;span class="line"&gt;&lt;span class="cl"&gt;touch index.js
&lt;/span&gt;&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;&lt;/div&gt;&lt;p&gt;This will create a folder for our express API backend and create a JS file to write code in. Edit the &lt;code&gt;index.js&lt;/code&gt; file and add this code:&lt;/p&gt;
&lt;div class="highlight"&gt;&lt;pre tabindex="0" class="chroma"&gt;&lt;code class="language-javascript" data-lang="javascript"&gt;&lt;span class="line"&gt;&lt;span class="cl"&gt;&lt;span class="kr"&gt;const&lt;/span&gt; &lt;span class="nx"&gt;express&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="nx"&gt;require&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="s1"&gt;&amp;#39;express&amp;#39;&lt;/span&gt;&lt;span class="p"&gt;);&lt;/span&gt;
&lt;/span&gt;&lt;/span&gt;&lt;span class="line"&gt;&lt;span class="cl"&gt;&lt;span class="kr"&gt;const&lt;/span&gt; &lt;span class="nx"&gt;app&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="nx"&gt;express&lt;/span&gt;&lt;span class="p"&gt;();&lt;/span&gt;
&lt;/span&gt;&lt;/span&gt;&lt;span class="line"&gt;&lt;span class="cl"&gt;
&lt;/span&gt;&lt;/span&gt;&lt;span class="line"&gt;&lt;span class="cl"&gt;&lt;span class="nx"&gt;app&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nx"&gt;get&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="s1"&gt;&amp;#39;/&amp;#39;&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="nx"&gt;req&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="nx"&gt;res&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt; &lt;span class="p"&gt;=&amp;gt;&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt;
&lt;/span&gt;&lt;/span&gt;&lt;span class="line"&gt;&lt;span class="cl"&gt; &lt;span class="nx"&gt;res&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nx"&gt;send&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="s1"&gt;&amp;#39;Hello World&amp;#39;&lt;/span&gt;&lt;span class="p"&gt;);&lt;/span&gt;
&lt;/span&gt;&lt;/span&gt;&lt;span class="line"&gt;&lt;span class="cl"&gt;&lt;span class="p"&gt;});&lt;/span&gt;
&lt;/span&gt;&lt;/span&gt;&lt;span class="line"&gt;&lt;span class="cl"&gt;
&lt;/span&gt;&lt;/span&gt;&lt;span class="line"&gt;&lt;span class="cl"&gt;&lt;span class="nx"&gt;app&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nx"&gt;listen&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="mi"&gt;3000&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="p"&gt;()&lt;/span&gt; &lt;span class="p"&gt;=&amp;gt;&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt;
&lt;/span&gt;&lt;/span&gt;&lt;span class="line"&gt;&lt;span class="cl"&gt; &lt;span class="nx"&gt;console&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nx"&gt;log&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="s1"&gt;&amp;#39;Server is running on port 3000&amp;#39;&lt;/span&gt;&lt;span class="p"&gt;);&lt;/span&gt;
&lt;/span&gt;&lt;/span&gt;&lt;span class="line"&gt;&lt;span class="cl"&gt;&lt;span class="p"&gt;});&lt;/span&gt;
&lt;/span&gt;&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;&lt;/div&gt;&lt;p&gt;We have just initialised an express backend which is serving the &amp;ldquo;Hello World&amp;rdquo; string at the root (/) endpoint. i.e. on sending a GET request to &lt;code&gt;/&lt;/code&gt;, &amp;ldquo;Hello World&amp;rdquo; will be returned.&lt;/p&gt;
&lt;p&gt;Run this file with:&lt;/p&gt;
&lt;div class="highlight"&gt;&lt;pre tabindex="0" class="chroma"&gt;&lt;code class="language-bash" data-lang="bash"&gt;&lt;span class="line"&gt;&lt;span class="cl"&gt;node index.js
&lt;/span&gt;&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;&lt;/div&gt;&lt;p&gt;And open &lt;code&gt;http://localhost:3000/&lt;/code&gt; on any browser. It should show &amp;ldquo;Hello World&amp;rdquo;.&lt;/p&gt;
&lt;figure class="doc-figure"&gt;&lt;img src="https://miro.medium.com/v2/resize:fit:1324/format:webp/1*fh2r1wAEpvcPK0bZpV9Y-Q.png" alt="Hello World Demo" loading="lazy" decoding="async"&gt;&lt;figcaption aria-hidden="true"&gt;Hello World Demo&lt;/figcaption&gt;&lt;/figure&gt;&lt;p&gt;Try changing the string in &lt;code&gt;res.send('this one')&lt;/code&gt; to simple HTML like &lt;code&gt;&amp;lt;h1&amp;gt;Hi there&amp;lt;/h1&amp;gt;&lt;/code&gt; and restart the express server (ctrl+c and &lt;code&gt;node index.js&lt;/code&gt; again), check if something changes on the served page.&lt;/p&gt;
&lt;h2 id="defining-a-post-endpoint"&gt;Defining a POST Endpoint&lt;a class="heading-anchor" href="#defining-a-post-endpoint" aria-label="Link to this section"&gt;&lt;/a&gt;
&lt;/h2&gt;
&lt;p&gt;But this is just for GET requests, we cannot make POST requests to this endpoint.&lt;/p&gt;
&lt;figure class="doc-figure"&gt;&lt;img src="https://miro.medium.com/v2/resize:fit:1236/format:webp/1*cJV5BQAHvocrDKE_LO1jMQ.png" alt="POST Request Error" loading="lazy" decoding="async"&gt;&lt;figcaption aria-hidden="true"&gt;POST Request Error&lt;/figcaption&gt;&lt;/figure&gt;&lt;p&gt;As you can see it gives an error if we try to make a POST request to the GET endpoint we created. Let&amp;rsquo;s define a POST endpoint now. Add this after defining the app:&lt;/p&gt;
&lt;div class="highlight"&gt;&lt;pre tabindex="0" class="chroma"&gt;&lt;code class="language-javascript" data-lang="javascript"&gt;&lt;span class="line"&gt;&lt;span class="cl"&gt;&lt;span class="nx"&gt;app&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nx"&gt;use&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="nx"&gt;express&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nx"&gt;json&lt;/span&gt;&lt;span class="p"&gt;());&lt;/span&gt;
&lt;/span&gt;&lt;/span&gt;&lt;span class="line"&gt;&lt;span class="cl"&gt;
&lt;/span&gt;&lt;/span&gt;&lt;span class="line"&gt;&lt;span class="cl"&gt;&lt;span class="c1"&gt;// POST endpoint
&lt;/span&gt;&lt;/span&gt;&lt;/span&gt;&lt;span class="line"&gt;&lt;span class="cl"&gt;&lt;span class="nx"&gt;app&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nx"&gt;post&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="s1"&gt;&amp;#39;/&amp;#39;&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="nx"&gt;req&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="nx"&gt;res&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt; &lt;span class="p"&gt;=&amp;gt;&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt;
&lt;/span&gt;&lt;/span&gt;&lt;span class="line"&gt;&lt;span class="cl"&gt; &lt;span class="kr"&gt;const&lt;/span&gt; &lt;span class="nx"&gt;name&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="nx"&gt;req&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nx"&gt;body&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nx"&gt;name&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt;
&lt;/span&gt;&lt;/span&gt;&lt;span class="line"&gt;&lt;span class="cl"&gt; &lt;span class="nx"&gt;res&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nx"&gt;send&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="sb"&gt;`Hello &lt;/span&gt;&lt;span class="si"&gt;${&lt;/span&gt;&lt;span class="nx"&gt;name&lt;/span&gt;&lt;span class="si"&gt;}&lt;/span&gt;&lt;span class="sb"&gt;!`&lt;/span&gt;&lt;span class="p"&gt;);&lt;/span&gt;
&lt;/span&gt;&lt;/span&gt;&lt;span class="line"&gt;&lt;span class="cl"&gt;&lt;span class="p"&gt;});&lt;/span&gt;
&lt;/span&gt;&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;&lt;/div&gt;&lt;p&gt;Restart the app and let&amp;rsquo;s try sending a post request to the server. You can use any HTTP client like Thunder Client on VSCode or Postman. I am going to use cURL coz simple:&lt;/p&gt;
&lt;div class="highlight"&gt;&lt;pre tabindex="0" class="chroma"&gt;&lt;code class="language-bash" data-lang="bash"&gt;&lt;span class="line"&gt;&lt;span class="cl"&gt;curl http://localhost:3000/ -d &lt;span class="s1"&gt;&amp;#39;{ &amp;#34;name&amp;#34;: &amp;#34;Ankush&amp;#34; }&amp;#39;&lt;/span&gt; -H &lt;span class="s2"&gt;&amp;#34;Content-Type: application/json&amp;#34;&lt;/span&gt;
&lt;/span&gt;&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;&lt;/div&gt;&lt;p&gt;This should print the response as:&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Hello Ankush!&lt;/strong&gt;&lt;/p&gt;
&lt;p&gt;BOOM! You just created a GET and POST request! A majority of web APIs are comprised of just GET and POST requests.&lt;/p&gt;
&lt;h2 id="time-to-write-some-frontend"&gt;Time to Write Some Frontend&lt;a class="heading-anchor" href="#time-to-write-some-frontend" aria-label="Link to this section"&gt;&lt;/a&gt;
&lt;/h2&gt;
&lt;p&gt;Now let&amp;rsquo;s serve an HTML page that interacts with our POST endpoint:&lt;/p&gt;
&lt;div class="highlight"&gt;&lt;pre tabindex="0" class="chroma"&gt;&lt;code class="language-html" data-lang="html"&gt;&lt;span class="line"&gt;&lt;span class="cl"&gt;&lt;span class="p"&gt;&amp;lt;&lt;/span&gt;&lt;span class="nt"&gt;html&lt;/span&gt;&lt;span class="p"&gt;&amp;gt;&lt;/span&gt;
&lt;/span&gt;&lt;/span&gt;&lt;span class="line"&gt;&lt;span class="cl"&gt; &lt;span class="p"&gt;&amp;lt;&lt;/span&gt;&lt;span class="nt"&gt;body&lt;/span&gt;&lt;span class="p"&gt;&amp;gt;&lt;/span&gt;
&lt;/span&gt;&lt;/span&gt;&lt;span class="line"&gt;&lt;span class="cl"&gt; &lt;span class="p"&gt;&amp;lt;&lt;/span&gt;&lt;span class="nt"&gt;input&lt;/span&gt; &lt;span class="na"&gt;type&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="s"&gt;&amp;#34;text&amp;#34;&lt;/span&gt; &lt;span class="na"&gt;id&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="s"&gt;&amp;#34;name&amp;#34;&lt;/span&gt; &lt;span class="na"&gt;name&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="s"&gt;&amp;#34;name&amp;#34;&lt;/span&gt; &lt;span class="na"&gt;placeholder&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="s"&gt;&amp;#34;Enter your name&amp;#34;&lt;/span&gt;&lt;span class="p"&gt;&amp;gt;&lt;/span&gt;
&lt;/span&gt;&lt;/span&gt;&lt;span class="line"&gt;&lt;span class="cl"&gt; &lt;span class="p"&gt;&amp;lt;&lt;/span&gt;&lt;span class="nt"&gt;button&lt;/span&gt; &lt;span class="na"&gt;onclick&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="s"&gt;&amp;#34;handleSubmit()&amp;#34;&lt;/span&gt;&lt;span class="p"&gt;&amp;gt;&lt;/span&gt;Submit&lt;span class="p"&gt;&amp;lt;/&lt;/span&gt;&lt;span class="nt"&gt;button&lt;/span&gt;&lt;span class="p"&gt;&amp;gt;&lt;/span&gt;
&lt;/span&gt;&lt;/span&gt;&lt;span class="line"&gt;&lt;span class="cl"&gt; &lt;span class="p"&gt;&amp;lt;&lt;/span&gt;&lt;span class="nt"&gt;script&lt;/span&gt;&lt;span class="p"&gt;&amp;gt;&lt;/span&gt;
&lt;/span&gt;&lt;/span&gt;&lt;span class="line"&gt;&lt;span class="cl"&gt; &lt;span class="kr"&gt;async&lt;/span&gt; &lt;span class="kd"&gt;function&lt;/span&gt; &lt;span class="nx"&gt;handleSubmit&lt;/span&gt;&lt;span class="p"&gt;()&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt;
&lt;/span&gt;&lt;/span&gt;&lt;span class="line"&gt;&lt;span class="cl"&gt; &lt;span class="kr"&gt;const&lt;/span&gt; &lt;span class="nx"&gt;name&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="nb"&gt;document&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nx"&gt;getElementById&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="s1"&gt;&amp;#39;name&amp;#39;&lt;/span&gt;&lt;span class="p"&gt;).&lt;/span&gt;&lt;span class="nx"&gt;value&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt;
&lt;/span&gt;&lt;/span&gt;&lt;span class="line"&gt;&lt;span class="cl"&gt; &lt;span class="kr"&gt;const&lt;/span&gt; &lt;span class="nx"&gt;response&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="kr"&gt;await&lt;/span&gt; &lt;span class="nx"&gt;fetch&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="s1"&gt;&amp;#39;/&amp;#39;&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt;
&lt;/span&gt;&lt;/span&gt;&lt;span class="line"&gt;&lt;span class="cl"&gt; &lt;span class="nx"&gt;method&lt;/span&gt;&lt;span class="o"&gt;:&lt;/span&gt; &lt;span class="s1"&gt;&amp;#39;POST&amp;#39;&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
&lt;/span&gt;&lt;/span&gt;&lt;span class="line"&gt;&lt;span class="cl"&gt; &lt;span class="nx"&gt;headers&lt;/span&gt;&lt;span class="o"&gt;:&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt;
&lt;/span&gt;&lt;/span&gt;&lt;span class="line"&gt;&lt;span class="cl"&gt; &lt;span class="s1"&gt;&amp;#39;Content-Type&amp;#39;&lt;/span&gt;&lt;span class="o"&gt;:&lt;/span&gt; &lt;span class="s1"&gt;&amp;#39;application/json&amp;#39;&lt;/span&gt;
&lt;/span&gt;&lt;/span&gt;&lt;span class="line"&gt;&lt;span class="cl"&gt; &lt;span class="p"&gt;},&lt;/span&gt;
&lt;/span&gt;&lt;/span&gt;&lt;span class="line"&gt;&lt;span class="cl"&gt; &lt;span class="nx"&gt;body&lt;/span&gt;&lt;span class="o"&gt;:&lt;/span&gt; &lt;span class="nx"&gt;JSON&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nx"&gt;stringify&lt;/span&gt;&lt;span class="p"&gt;({&lt;/span&gt; &lt;span class="nx"&gt;name&lt;/span&gt; &lt;span class="p"&gt;})&lt;/span&gt;
&lt;/span&gt;&lt;/span&gt;&lt;span class="line"&gt;&lt;span class="cl"&gt; &lt;span class="p"&gt;});&lt;/span&gt;
&lt;/span&gt;&lt;/span&gt;&lt;span class="line"&gt;&lt;span class="cl"&gt; &lt;span class="kr"&gt;const&lt;/span&gt; &lt;span class="nx"&gt;data&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="kr"&gt;await&lt;/span&gt; &lt;span class="nx"&gt;response&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nx"&gt;text&lt;/span&gt;&lt;span class="p"&gt;();&lt;/span&gt;
&lt;/span&gt;&lt;/span&gt;&lt;span class="line"&gt;&lt;span class="cl"&gt; &lt;span class="nb"&gt;document&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nx"&gt;body&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nx"&gt;innerHTML&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="nx"&gt;data&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt;
&lt;/span&gt;&lt;/span&gt;&lt;span class="line"&gt;&lt;span class="cl"&gt; &lt;span class="p"&gt;}&lt;/span&gt;
&lt;/span&gt;&lt;/span&gt;&lt;span class="line"&gt;&lt;span class="cl"&gt; &lt;span class="p"&gt;&amp;lt;/&lt;/span&gt;&lt;span class="nt"&gt;script&lt;/span&gt;&lt;span class="p"&gt;&amp;gt;&lt;/span&gt;
&lt;/span&gt;&lt;/span&gt;&lt;span class="line"&gt;&lt;span class="cl"&gt; &lt;span class="p"&gt;&amp;lt;/&lt;/span&gt;&lt;span class="nt"&gt;body&lt;/span&gt;&lt;span class="p"&gt;&amp;gt;&lt;/span&gt;
&lt;/span&gt;&lt;/span&gt;&lt;span class="line"&gt;&lt;span class="cl"&gt;&lt;span class="p"&gt;&amp;lt;/&lt;/span&gt;&lt;span class="nt"&gt;html&lt;/span&gt;&lt;span class="p"&gt;&amp;gt;&lt;/span&gt;
&lt;/span&gt;&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;&lt;/div&gt;&lt;p&gt;We will serve this HTML through the GET endpoint we created earlier:&lt;/p&gt;
&lt;div class="highlight"&gt;&lt;pre tabindex="0" class="chroma"&gt;&lt;code class="language-javascript" data-lang="javascript"&gt;&lt;span class="line"&gt;&lt;span class="cl"&gt;&lt;span class="nx"&gt;app&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nx"&gt;get&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="s1"&gt;&amp;#39;/&amp;#39;&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="nx"&gt;req&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="nx"&gt;res&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt; &lt;span class="p"&gt;=&amp;gt;&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt;
&lt;/span&gt;&lt;/span&gt;&lt;span class="line"&gt;&lt;span class="cl"&gt; &lt;span class="nx"&gt;res&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nx"&gt;send&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="sb"&gt;`
&lt;/span&gt;&lt;/span&gt;&lt;/span&gt;&lt;span class="line"&gt;&lt;span class="cl"&gt;&lt;span class="sb"&gt; &amp;lt;html&amp;gt;
&lt;/span&gt;&lt;/span&gt;&lt;/span&gt;&lt;span class="line"&gt;&lt;span class="cl"&gt;&lt;span class="sb"&gt; &amp;lt;body&amp;gt;
&lt;/span&gt;&lt;/span&gt;&lt;/span&gt;&lt;span class="line"&gt;&lt;span class="cl"&gt;&lt;span class="sb"&gt; &amp;lt;input type=&amp;#34;text&amp;#34; id=&amp;#34;name&amp;#34; name=&amp;#34;name&amp;#34; placeholder=&amp;#34;Enter your name&amp;#34;&amp;gt;
&lt;/span&gt;&lt;/span&gt;&lt;/span&gt;&lt;span class="line"&gt;&lt;span class="cl"&gt;&lt;span class="sb"&gt; &amp;lt;button onclick=&amp;#34;handleSubmit()&amp;#34;&amp;gt;Submit&amp;lt;/button&amp;gt;
&lt;/span&gt;&lt;/span&gt;&lt;/span&gt;&lt;span class="line"&gt;&lt;span class="cl"&gt;&lt;span class="sb"&gt; &amp;lt;script&amp;gt;
&lt;/span&gt;&lt;/span&gt;&lt;/span&gt;&lt;span class="line"&gt;&lt;span class="cl"&gt;&lt;span class="sb"&gt; async function handleSubmit() {
&lt;/span&gt;&lt;/span&gt;&lt;/span&gt;&lt;span class="line"&gt;&lt;span class="cl"&gt;&lt;span class="sb"&gt; const name = document.getElementById(&amp;#39;name&amp;#39;).value;
&lt;/span&gt;&lt;/span&gt;&lt;/span&gt;&lt;span class="line"&gt;&lt;span class="cl"&gt;&lt;span class="sb"&gt; const response = await fetch(&amp;#39;/&amp;#39;, {
&lt;/span&gt;&lt;/span&gt;&lt;/span&gt;&lt;span class="line"&gt;&lt;span class="cl"&gt;&lt;span class="sb"&gt; method: &amp;#39;POST&amp;#39;,
&lt;/span&gt;&lt;/span&gt;&lt;/span&gt;&lt;span class="line"&gt;&lt;span class="cl"&gt;&lt;span class="sb"&gt; headers: {
&lt;/span&gt;&lt;/span&gt;&lt;/span&gt;&lt;span class="line"&gt;&lt;span class="cl"&gt;&lt;span class="sb"&gt; &amp;#39;Content-Type&amp;#39;: &amp;#39;application/json&amp;#39;
&lt;/span&gt;&lt;/span&gt;&lt;/span&gt;&lt;span class="line"&gt;&lt;span class="cl"&gt;&lt;span class="sb"&gt; },
&lt;/span&gt;&lt;/span&gt;&lt;/span&gt;&lt;span class="line"&gt;&lt;span class="cl"&gt;&lt;span class="sb"&gt; body: JSON.stringify({ name })
&lt;/span&gt;&lt;/span&gt;&lt;/span&gt;&lt;span class="line"&gt;&lt;span class="cl"&gt;&lt;span class="sb"&gt; });
&lt;/span&gt;&lt;/span&gt;&lt;/span&gt;&lt;span class="line"&gt;&lt;span class="cl"&gt;&lt;span class="sb"&gt; const data = await response.text();
&lt;/span&gt;&lt;/span&gt;&lt;/span&gt;&lt;span class="line"&gt;&lt;span class="cl"&gt;&lt;span class="sb"&gt; document.body.innerHTML = data;
&lt;/span&gt;&lt;/span&gt;&lt;/span&gt;&lt;span class="line"&gt;&lt;span class="cl"&gt;&lt;span class="sb"&gt; }
&lt;/span&gt;&lt;/span&gt;&lt;/span&gt;&lt;span class="line"&gt;&lt;span class="cl"&gt;&lt;span class="sb"&gt; &amp;lt;/script&amp;gt;
&lt;/span&gt;&lt;/span&gt;&lt;/span&gt;&lt;span class="line"&gt;&lt;span class="cl"&gt;&lt;span class="sb"&gt; &amp;lt;/body&amp;gt;
&lt;/span&gt;&lt;/span&gt;&lt;/span&gt;&lt;span class="line"&gt;&lt;span class="cl"&gt;&lt;span class="sb"&gt; &amp;lt;/html&amp;gt;
&lt;/span&gt;&lt;/span&gt;&lt;/span&gt;&lt;span class="line"&gt;&lt;span class="cl"&gt;&lt;span class="sb"&gt; `&lt;/span&gt;&lt;span class="p"&gt;);&lt;/span&gt;
&lt;/span&gt;&lt;/span&gt;&lt;span class="line"&gt;&lt;span class="cl"&gt;&lt;span class="p"&gt;});&lt;/span&gt;
&lt;/span&gt;&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;&lt;/div&gt;&lt;p&gt;Restart the express app and refresh &lt;code&gt;localhost:3000&lt;/code&gt;, and try entering your name.&lt;/p&gt;
&lt;figure class="doc-figure"&gt;&lt;img src="https://miro.medium.com/v2/resize:fit:1400/format:webp/1*KhXDm_zbhiECJnoVBZE-XA.png" alt="Frontend Demo with Network Tab" loading="lazy" decoding="async"&gt;&lt;figcaption aria-hidden="true"&gt;Frontend Demo with Network Tab&lt;/figcaption&gt;&lt;/figure&gt;&lt;p&gt;Now if you open developer tools and go into the network tab, you can see the POST request being sent and the different data associated with it.&lt;/p&gt;
&lt;figure class="doc-figure"&gt;&lt;img src="https://miro.medium.com/v2/resize:fit:1400/format:webp/1*Wkbk2CrKOqYMwyQ11P9Eiw.png" alt="Network Tab" loading="lazy" decoding="async"&gt;&lt;figcaption aria-hidden="true"&gt;Network Tab&lt;/figcaption&gt;&lt;/figure&gt;&lt;hr&gt;
&lt;p&gt;And that&amp;rsquo;s a wrap! You&amp;rsquo;ve just learned the basics of web APIs, REST principles, and created your own simple API with both GET and POST endpoints. Now go build something cool! 🚀&lt;/p&gt;</content:encoded></item><item><title>Automating Mandelbrot Fractal generation with SYCL Programming</title><link>https://ankush.one/blogs/mandelbrot-intel-oneapi/</link><pubDate>Thu, 30 Mar 2023 00:00:00 +0000</pubDate><author>ankush4singh@gmail.com (Ankush Singh)</author><guid>https://ankush.one/blogs/mandelbrot-intel-oneapi/</guid><category>graphics</category><category>hackathon</category><category>tutorial</category><description>A walkthrough of generating Mandelbrot fractals with Intel oneAPI SYCL, plus an automation script — from a hackathon-winning project at IIT Roorkee.</description><content:encoded>&lt;h2 id="whats-this-about"&gt;What’s this about?&lt;a class="heading-anchor" href="#whats-this-about" aria-label="Link to this section"&gt;&lt;/a&gt;
&lt;/h2&gt;
&lt;p&gt;I had recently attended a workshop by Intel on their oneAPI programming model which was conducted by &lt;a href="https://medium.com/u/af7b551ddcdb" data-external&gt;Abhishek Nandy&lt;/a&gt; at IIT Roorkee during the Cognizance 2023. I attended this workshop with my friend, &lt;a href="https://medium.com/@kragrrr" data-external&gt;Krish Agrawal&lt;/a&gt;, who is also the co-author of this blog.&lt;br&gt;
The workshop was held for 2 days in which Nandy Sir on the first day told about this new model by &lt;strong&gt;Intel&lt;/strong&gt; and how to get started with its implementation with the help of &lt;strong&gt;JupyterLab&lt;/strong&gt;. Before ending the day, we were told that there will be a hackathon where he will take implementations of the this new model by all the students in the form of teams as to how we use Intel oneAPI in real-life projects.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;SYCL (/sɪkl/)&lt;/strong&gt; is a programming model for high-performance computing that allows developers to write code for heterogeneous systems that use accelerators such as GPUs, FPGAs, and other specialised processing units. It is an open standard developed by the Khronos Group, an industry consortium that also develops other widely used graphics and compute APIs such as OpenGL and Vulkan. SYCL code can be executed on a variety of devices, including CPUs, GPUs, and FPGAs, without modification.&lt;/p&gt;
&lt;figure class="doc-figure"&gt;&lt;img src="https://miro.medium.com/v2/resize:fit:1400/format:webp/1*HiYIE_ehe_7jpMEYs9meew.jpeg" alt="Mandelbrot fractal" loading="lazy" decoding="async"&gt;&lt;figcaption aria-hidden="true"&gt;Mandelbrot fractal&lt;/figcaption&gt;&lt;/figure&gt;&lt;p&gt;&lt;strong&gt;Fractals&lt;/strong&gt; are fascinating geometric shapes that repeat themselves infinitely on different scales. They are created through a process of &lt;strong&gt;repeating a simple mathematical equation&lt;/strong&gt; or algorithm to form increasingly intricate patterns. These fractal shapes are naturally occurring and can be observed in many aspects of nature, such as the patterns of a fern leaf, the coastlines of continents, and even the formation of clouds. One of the most well-known and captivating fractals is the &lt;strong&gt;Mandelbrot fractal&lt;/strong&gt;. This type of fractal is generated through an iterative process using complex numbers, resulting in a shape that is renowned for its intricate detail and infinite complexity. It’s a true marvel of mathematics!&lt;/p&gt;
&lt;p&gt;Now that we know what our project revolves around, let’s see its implementation in JupyterLab.&lt;/p&gt;
&lt;h2 id="getting-started-on-devcloud"&gt;Getting started on DevCloud&lt;a class="heading-anchor" href="#getting-started-on-devcloud" aria-label="Link to this section"&gt;&lt;/a&gt;
&lt;/h2&gt;
&lt;p&gt;DevCloud is sandbox that intel provides to learn and try out their oneAPI ecosystem. We will be using the devcloud jupyter lab to run our SYCL fractal generation project. It is way easier and quicker to run SYCL projects on the sandbox than on your local machine.&lt;/p&gt;
&lt;p&gt;So head over to &lt;a href="https://web.archive.org/web/20241207195212/https://devcloud.intel.com/oneapi/" data-external&gt;devcloud.intel.com/oneapi&lt;/a&gt; (archived snapshot, the DevCloud site is no longer online) and create an account. Navigate to the &amp;lsquo;Getting Started&amp;rsquo; tab, scroll down to &amp;lsquo;Connect with JupyterLab&amp;rsquo; and Launch a JupyterLab instance. It might take some time or get loaded instantly depending on the server load so have some patience.&lt;/p&gt;
&lt;figure class="doc-figure"&gt;&lt;img src="https://miro.medium.com/v2/resize:fit:1400/format:webp/1*8JUHbOpbnWCiTXeLqpqEPQ.png" alt="Devcloud Jupyterlab" loading="lazy" decoding="async"&gt;&lt;figcaption aria-hidden="true"&gt;Devcloud Jupyterlab&lt;/figcaption&gt;&lt;/figure&gt;&lt;p&gt;Once launched, you will find a bunch of folders in the file explorer, including a &lt;code&gt;mandelbrot&lt;/code&gt; folder.&lt;/p&gt;
&lt;p&gt;Open the Terminal from the landing page, cd into the &lt;code&gt;mandelbrot&lt;/code&gt; directory and run the build script.&lt;/p&gt;
&lt;div class="highlight"&gt;&lt;pre tabindex="0" class="chroma"&gt;&lt;code class="language-bash" data-lang="bash"&gt;&lt;span class="line"&gt;&lt;span class="cl"&gt;&lt;span class="nb"&gt;cd&lt;/span&gt; mandelbrot
&lt;/span&gt;&lt;/span&gt;&lt;span class="line"&gt;&lt;span class="cl"&gt;chmod +x build.sh
&lt;/span&gt;&lt;/span&gt;&lt;span class="line"&gt;&lt;span class="cl"&gt;./build.sh
&lt;/span&gt;&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;&lt;/div&gt;&lt;figure class="doc-figure"&gt;&lt;img src="https://miro.medium.com/v2/resize:fit:1400/format:webp/1*62IvxnuK0M2Maf3cG43Cjw.png" alt="Build results" loading="lazy" decoding="async"&gt;&lt;figcaption aria-hidden="true"&gt;Build results&lt;/figcaption&gt;&lt;/figure&gt;&lt;p&gt;This will compile and run the project and create a &lt;code&gt;mandelbrot.png&lt;/code&gt; file inside the &lt;code&gt;build&lt;/code&gt; folder. As you can see the image is 1024x1024 pixels and it took just 190 milliseconds to generate the image (serial time + parallel time). If we were to do the same in python or regular C++ without SYCL, it would take seconds to generate the same result. So that’s the power of SYCL.&lt;/p&gt;
&lt;p&gt;If you checkout the &lt;code&gt;src/mandel.hpp&lt;/code&gt; file, you can find 4 variables row and column size, max iterations and repetitions. We can modify these values to generate Mandelbrot sets of different resolutions and qualities.&lt;/p&gt;
&lt;p&gt;Try changing these values and see how much time it takes for the images of different configurations to generate.&lt;/p&gt;
&lt;h2 id="automate-it"&gt;Automate it&lt;a class="heading-anchor" href="#automate-it" aria-label="Link to this section"&gt;&lt;/a&gt;
&lt;/h2&gt;
&lt;p&gt;Let’s now write a &lt;strong&gt;bash script&lt;/strong&gt; that automatically changes the values in the C++ file and builds the project.&lt;/p&gt;
&lt;p&gt;Create a file &lt;code&gt;autogen.sh&lt;/code&gt; and inside it write the following lines.&lt;/p&gt;
&lt;div class="highlight"&gt;&lt;pre tabindex="0" class="chroma"&gt;&lt;code class="language-bash" data-lang="bash"&gt;&lt;span class="line"&gt;&lt;span class="cl"&gt;&lt;span class="cp"&gt;#!/bin/bash
&lt;/span&gt;&lt;/span&gt;&lt;/span&gt;&lt;span class="line"&gt;&lt;span class="cl"&gt;
&lt;/span&gt;&lt;/span&gt;&lt;span class="line"&gt;&lt;span class="cl"&gt;&lt;span class="nv"&gt;PASSES&lt;/span&gt;&lt;span class="o"&gt;=(&lt;/span&gt;&lt;span class="m"&gt;10&lt;/span&gt; &lt;span class="m"&gt;100&lt;/span&gt; &lt;span class="m"&gt;500&lt;/span&gt; &lt;span class="m"&gt;1000&lt;/span&gt; 5000&lt;span class="o"&gt;)&lt;/span&gt;
&lt;/span&gt;&lt;/span&gt;&lt;span class="line"&gt;&lt;span class="cl"&gt;&lt;span class="nv"&gt;SIZES&lt;/span&gt;&lt;span class="o"&gt;=(&lt;/span&gt;&lt;span class="m"&gt;1024&lt;/span&gt; &lt;span class="m"&gt;2048&lt;/span&gt; &lt;span class="m"&gt;4096&lt;/span&gt; &lt;span class="m"&gt;8192&lt;/span&gt; 16384&lt;span class="o"&gt;)&lt;/span&gt;
&lt;/span&gt;&lt;/span&gt;&lt;span class="line"&gt;&lt;span class="cl"&gt;
&lt;/span&gt;&lt;/span&gt;&lt;span class="line"&gt;&lt;span class="cl"&gt;&lt;span class="nv"&gt;EDIT_FILE&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="s2"&gt;&amp;#34;src/mandel.hpp&amp;#34;&lt;/span&gt;
&lt;/span&gt;&lt;/span&gt;&lt;span class="line"&gt;&lt;span class="cl"&gt;&lt;span class="nv"&gt;BUILD_FILE&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="s2"&gt;&amp;#34;build.sh&amp;#34;&lt;/span&gt;
&lt;/span&gt;&lt;/span&gt;&lt;span class="line"&gt;&lt;span class="cl"&gt;
&lt;/span&gt;&lt;/span&gt;&lt;span class="line"&gt;&lt;span class="cl"&gt;&lt;span class="nv"&gt;prev_size&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="m"&gt;100&lt;/span&gt;
&lt;/span&gt;&lt;/span&gt;&lt;span class="line"&gt;&lt;span class="cl"&gt;&lt;span class="nv"&gt;prev_pass&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="m"&gt;100&lt;/span&gt;
&lt;/span&gt;&lt;/span&gt;&lt;span class="line"&gt;&lt;span class="cl"&gt;&lt;span class="nb"&gt;let&lt;/span&gt; &lt;span class="nv"&gt;count&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="m"&gt;0&lt;/span&gt;
&lt;/span&gt;&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;&lt;/div&gt;&lt;p&gt;The variable &lt;code&gt;PASSES&lt;/code&gt; contains a list of iterations that will be used in each image of the image generated and sizes contains the list of resolutions for the images in pixels. &lt;code&gt;prev_size&lt;/code&gt; and &lt;code&gt;prev_pass&lt;/code&gt; is a placeholder that we will use later. &lt;code&gt;count&lt;/code&gt; is the number of images generated.&lt;/p&gt;
&lt;div class="highlight"&gt;&lt;pre tabindex="0" class="chroma"&gt;&lt;code class="language-bash" data-lang="bash"&gt;&lt;span class="line"&gt;&lt;span class="cl"&gt;&lt;span class="k"&gt;for&lt;/span&gt; SIZE in &lt;span class="si"&gt;${&lt;/span&gt;&lt;span class="nv"&gt;SIZES&lt;/span&gt;&lt;span class="p"&gt;[@]&lt;/span&gt;&lt;span class="si"&gt;}&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt; &lt;span class="k"&gt;do&lt;/span&gt;
&lt;/span&gt;&lt;/span&gt;&lt;span class="line"&gt;&lt;span class="cl"&gt; &lt;span class="k"&gt;for&lt;/span&gt; PASS in &lt;span class="si"&gt;${&lt;/span&gt;&lt;span class="nv"&gt;PASSES&lt;/span&gt;&lt;span class="p"&gt;[@]&lt;/span&gt;&lt;span class="si"&gt;}&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt; &lt;span class="k"&gt;do&lt;/span&gt;
&lt;/span&gt;&lt;/span&gt;&lt;span class="line"&gt;&lt;span class="cl"&gt; &lt;span class="nb"&gt;echo&lt;/span&gt; &lt;span class="s2"&gt;&amp;#34;&amp;#34;&lt;/span&gt;
&lt;/span&gt;&lt;/span&gt;&lt;span class="line"&gt;&lt;span class="cl"&gt; &lt;span class="nb"&gt;echo&lt;/span&gt; &lt;span class="s2"&gt;&amp;#34;Generating fractal for &lt;/span&gt;&lt;span class="nv"&gt;$SIZE&lt;/span&gt;&lt;span class="s2"&gt; x &lt;/span&gt;&lt;span class="nv"&gt;$SIZE&lt;/span&gt;&lt;span class="s2"&gt; with &lt;/span&gt;&lt;span class="nv"&gt;$PASS&lt;/span&gt;&lt;span class="s2"&gt; passes&amp;#34;&lt;/span&gt;
&lt;/span&gt;&lt;/span&gt;&lt;span class="line"&gt;&lt;span class="cl"&gt; &lt;span class="nb"&gt;echo&lt;/span&gt; &lt;span class="s2"&gt;&amp;#34;&amp;#34;&lt;/span&gt;
&lt;/span&gt;&lt;/span&gt;&lt;span class="line"&gt;&lt;span class="cl"&gt;
&lt;/span&gt;&lt;/span&gt;&lt;span class="line"&gt;&lt;span class="cl"&gt; ...
&lt;/span&gt;&lt;/span&gt;&lt;span class="line"&gt;&lt;span class="cl"&gt;
&lt;/span&gt;&lt;/span&gt;&lt;span class="line"&gt;&lt;span class="cl"&gt; &lt;span class="k"&gt;done&lt;/span&gt;
&lt;/span&gt;&lt;/span&gt;&lt;span class="line"&gt;&lt;span class="cl"&gt;&lt;span class="k"&gt;done&lt;/span&gt;
&lt;/span&gt;&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;&lt;/div&gt;&lt;p&gt;We create a nested loop for the values of &lt;code&gt;SIZES&lt;/code&gt; and &lt;code&gt;PASSES&lt;/code&gt;, so we can have all the combinations of sizes and passes to generate the images with. In the body of the loop we print a message to display to the user.&lt;/p&gt;
&lt;p&gt;In order to edit the values of the 3 variables in the &lt;code&gt;mandel.hpp&lt;/code&gt; file we will use the &lt;code&gt;sed&lt;/code&gt; command.&lt;/p&gt;
&lt;div class="highlight"&gt;&lt;pre tabindex="0" class="chroma"&gt;&lt;code class="language-bash" data-lang="bash"&gt;&lt;span class="line"&gt;&lt;span class="cl"&gt;sed -ie &lt;span class="s2"&gt;&amp;#34;s/row_size=&lt;/span&gt;&lt;span class="nv"&gt;$prev_size&lt;/span&gt;&lt;span class="s2"&gt;/row_size=&lt;/span&gt;&lt;span class="nv"&gt;$SIZE&lt;/span&gt;&lt;span class="s2"&gt;/g&amp;#34;&lt;/span&gt; &lt;span class="nv"&gt;$EDIT_FILE&lt;/span&gt;
&lt;/span&gt;&lt;/span&gt;&lt;span class="line"&gt;&lt;span class="cl"&gt;sed -ie &lt;span class="s2"&gt;&amp;#34;s/col_size=&lt;/span&gt;&lt;span class="nv"&gt;$prev_size&lt;/span&gt;&lt;span class="s2"&gt;/col_size=&lt;/span&gt;&lt;span class="nv"&gt;$SIZE&lt;/span&gt;&lt;span class="s2"&gt;/g&amp;#34;&lt;/span&gt; &lt;span class="nv"&gt;$EDIT_FILE&lt;/span&gt;
&lt;/span&gt;&lt;/span&gt;&lt;span class="line"&gt;&lt;span class="cl"&gt;sed -ie &lt;span class="s2"&gt;&amp;#34;s/max_iterations=&lt;/span&gt;&lt;span class="nv"&gt;$prev_pass&lt;/span&gt;&lt;span class="s2"&gt;/max_iterations=&lt;/span&gt;&lt;span class="nv"&gt;$PASS&lt;/span&gt;&lt;span class="s2"&gt;/g&amp;#34;&lt;/span&gt; &lt;span class="nv"&gt;$EDIT_FILE&lt;/span&gt;
&lt;/span&gt;&lt;/span&gt;&lt;span class="line"&gt;&lt;span class="cl"&gt;sed -ie &lt;span class="s2"&gt;&amp;#34;s/repetitions=&lt;/span&gt;&lt;span class="nv"&gt;$prev_pass&lt;/span&gt;&lt;span class="s2"&gt;/repetitions=&lt;/span&gt;&lt;span class="nv"&gt;$PASS&lt;/span&gt;&lt;span class="s2"&gt;/g&amp;#34;&lt;/span&gt; &lt;span class="nv"&gt;$EDIT_FILE&lt;/span&gt;
&lt;/span&gt;&lt;/span&gt;&lt;span class="line"&gt;&lt;span class="cl"&gt;
&lt;/span&gt;&lt;/span&gt;&lt;span class="line"&gt;&lt;span class="cl"&gt;&lt;span class="nv"&gt;prev_size&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="nv"&gt;$SIZE&lt;/span&gt;
&lt;/span&gt;&lt;/span&gt;&lt;span class="line"&gt;&lt;span class="cl"&gt;&lt;span class="nv"&gt;prev_pass&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="nv"&gt;$PASS&lt;/span&gt;
&lt;/span&gt;&lt;/span&gt;&lt;span class="line"&gt;&lt;span class="cl"&gt;&lt;span class="nv"&gt;count&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="k"&gt;$((&lt;/span&gt;count+1&lt;span class="k"&gt;))&lt;/span&gt;
&lt;/span&gt;&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;&lt;/div&gt;&lt;p&gt;This uses find and replace to replace the previous values of the variables with new values and updates the count in each iteration of the bash loop.&lt;/p&gt;
&lt;p&gt;To run the build command we simply do&lt;/p&gt;
&lt;div class="highlight"&gt;&lt;pre tabindex="0" class="chroma"&gt;&lt;code class="language-bash" data-lang="bash"&gt;&lt;span class="line"&gt;&lt;span class="cl"&gt;./&lt;span class="nv"&gt;$BUILD_FILE&lt;/span&gt;
&lt;/span&gt;&lt;/span&gt;&lt;span class="line"&gt;&lt;span class="cl"&gt;
&lt;/span&gt;&lt;/span&gt;&lt;span class="line"&gt;&lt;span class="cl"&gt;&lt;span class="nb"&gt;echo&lt;/span&gt; &lt;span class="s2"&gt;&amp;#34;&amp;#34;&lt;/span&gt;
&lt;/span&gt;&lt;/span&gt;&lt;span class="line"&gt;&lt;span class="cl"&gt;&lt;span class="nb"&gt;echo&lt;/span&gt; &lt;span class="s2"&gt;&amp;#34;Generated for &lt;/span&gt;&lt;span class="nv"&gt;$SIZE&lt;/span&gt;&lt;span class="s2"&gt; x &lt;/span&gt;&lt;span class="nv"&gt;$SIZE&lt;/span&gt;&lt;span class="s2"&gt; with &lt;/span&gt;&lt;span class="nv"&gt;$PASS&lt;/span&gt;&lt;span class="s2"&gt; passes&amp;#34;&lt;/span&gt;
&lt;/span&gt;&lt;/span&gt;&lt;span class="line"&gt;&lt;span class="cl"&gt;&lt;span class="nb"&gt;echo&lt;/span&gt; &lt;span class="s2"&gt;&amp;#34;&amp;#34;&lt;/span&gt;
&lt;/span&gt;&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;&lt;/div&gt;&lt;p&gt;The project is built and image is saved in the build folder, so we will have to copy it from there into a location of our choice. Create a folder images in the mandelbrot folder and add the following copy command to the &lt;code&gt;autogen.sh&lt;/code&gt; script&lt;/p&gt;
&lt;div class="highlight"&gt;&lt;pre tabindex="0" class="chroma"&gt;&lt;code class="language-bash" data-lang="bash"&gt;&lt;span class="line"&gt;&lt;span class="cl"&gt;cp &lt;span class="s2"&gt;&amp;#34;./build/mandelbrot.png&amp;#34;&lt;/span&gt; &lt;span class="s2"&gt;&amp;#34;./images/&lt;/span&gt;&lt;span class="nv"&gt;$count&lt;/span&gt;&lt;span class="s2"&gt;.mandelbrot_&lt;/span&gt;&lt;span class="si"&gt;${&lt;/span&gt;&lt;span class="nv"&gt;SIZE&lt;/span&gt;&lt;span class="si"&gt;}&lt;/span&gt;&lt;span class="s2"&gt;_&lt;/span&gt;&lt;span class="si"&gt;${&lt;/span&gt;&lt;span class="nv"&gt;PASS&lt;/span&gt;&lt;span class="si"&gt;}&lt;/span&gt;&lt;span class="s2"&gt;.png&amp;#34;&lt;/span&gt;
&lt;/span&gt;&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;&lt;/div&gt;&lt;p&gt;This copies the image from build into images folder and also renames it to contain useful information such as size and number of passes and the count.&lt;/p&gt;
&lt;p&gt;Run the autogen script and let the fractals be created!&lt;/p&gt;
&lt;div class="highlight"&gt;&lt;pre tabindex="0" class="chroma"&gt;&lt;code class="language-bash" data-lang="bash"&gt;&lt;span class="line"&gt;&lt;span class="cl"&gt;chmod +x autogen.sh
&lt;/span&gt;&lt;/span&gt;&lt;span class="line"&gt;&lt;span class="cl"&gt;./autogen.sh
&lt;/span&gt;&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;&lt;/div&gt;&lt;p&gt;After the script is done, you will find the images in the images folder.&lt;/p&gt;
&lt;p&gt;Here is the full autogen script&lt;/p&gt;
&lt;div class="highlight"&gt;&lt;pre tabindex="0" class="chroma"&gt;&lt;code class="language-bash" data-lang="bash"&gt;&lt;span class="line"&gt;&lt;span class="cl"&gt;&lt;span class="cp"&gt;#!/bin/bash
&lt;/span&gt;&lt;/span&gt;&lt;/span&gt;&lt;span class="line"&gt;&lt;span class="cl"&gt;
&lt;/span&gt;&lt;/span&gt;&lt;span class="line"&gt;&lt;span class="cl"&gt;&lt;span class="nv"&gt;PASSES&lt;/span&gt;&lt;span class="o"&gt;=(&lt;/span&gt;&lt;span class="m"&gt;10&lt;/span&gt; &lt;span class="m"&gt;100&lt;/span&gt; &lt;span class="m"&gt;500&lt;/span&gt; &lt;span class="m"&gt;1000&lt;/span&gt; 5000&lt;span class="o"&gt;)&lt;/span&gt;
&lt;/span&gt;&lt;/span&gt;&lt;span class="line"&gt;&lt;span class="cl"&gt;&lt;span class="nv"&gt;SIZES&lt;/span&gt;&lt;span class="o"&gt;=(&lt;/span&gt;&lt;span class="m"&gt;1024&lt;/span&gt; &lt;span class="m"&gt;2048&lt;/span&gt; &lt;span class="m"&gt;4096&lt;/span&gt; &lt;span class="m"&gt;8192&lt;/span&gt; 16384&lt;span class="o"&gt;)&lt;/span&gt;
&lt;/span&gt;&lt;/span&gt;&lt;span class="line"&gt;&lt;span class="cl"&gt;
&lt;/span&gt;&lt;/span&gt;&lt;span class="line"&gt;&lt;span class="cl"&gt;&lt;span class="nv"&gt;EDIT_FILE&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="s2"&gt;&amp;#34;src/mandel.hpp&amp;#34;&lt;/span&gt;
&lt;/span&gt;&lt;/span&gt;&lt;span class="line"&gt;&lt;span class="cl"&gt;&lt;span class="nv"&gt;BUILD_FILE&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="s2"&gt;&amp;#34;build.sh&amp;#34;&lt;/span&gt;
&lt;/span&gt;&lt;/span&gt;&lt;span class="line"&gt;&lt;span class="cl"&gt;
&lt;/span&gt;&lt;/span&gt;&lt;span class="line"&gt;&lt;span class="cl"&gt;&lt;span class="nv"&gt;prev_size&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="m"&gt;100&lt;/span&gt;
&lt;/span&gt;&lt;/span&gt;&lt;span class="line"&gt;&lt;span class="cl"&gt;&lt;span class="nv"&gt;prev_pass&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="m"&gt;100&lt;/span&gt;
&lt;/span&gt;&lt;/span&gt;&lt;span class="line"&gt;&lt;span class="cl"&gt;&lt;span class="nb"&gt;let&lt;/span&gt; &lt;span class="nv"&gt;count&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="m"&gt;0&lt;/span&gt;
&lt;/span&gt;&lt;/span&gt;&lt;span class="line"&gt;&lt;span class="cl"&gt;
&lt;/span&gt;&lt;/span&gt;&lt;span class="line"&gt;&lt;span class="cl"&gt;&lt;span class="k"&gt;for&lt;/span&gt; SIZE in &lt;span class="si"&gt;${&lt;/span&gt;&lt;span class="nv"&gt;SIZES&lt;/span&gt;&lt;span class="p"&gt;[@]&lt;/span&gt;&lt;span class="si"&gt;}&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt; &lt;span class="k"&gt;do&lt;/span&gt;
&lt;/span&gt;&lt;/span&gt;&lt;span class="line"&gt;&lt;span class="cl"&gt; &lt;span class="k"&gt;for&lt;/span&gt; PASS in &lt;span class="si"&gt;${&lt;/span&gt;&lt;span class="nv"&gt;PASSES&lt;/span&gt;&lt;span class="p"&gt;[@]&lt;/span&gt;&lt;span class="si"&gt;}&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt; &lt;span class="k"&gt;do&lt;/span&gt;
&lt;/span&gt;&lt;/span&gt;&lt;span class="line"&gt;&lt;span class="cl"&gt; &lt;span class="nb"&gt;echo&lt;/span&gt; &lt;span class="s2"&gt;&amp;#34;&amp;#34;&lt;/span&gt;
&lt;/span&gt;&lt;/span&gt;&lt;span class="line"&gt;&lt;span class="cl"&gt; &lt;span class="nb"&gt;echo&lt;/span&gt; &lt;span class="s2"&gt;&amp;#34;Generating fractal for &lt;/span&gt;&lt;span class="nv"&gt;$SIZE&lt;/span&gt;&lt;span class="s2"&gt; x &lt;/span&gt;&lt;span class="nv"&gt;$SIZE&lt;/span&gt;&lt;span class="s2"&gt; with &lt;/span&gt;&lt;span class="nv"&gt;$PASS&lt;/span&gt;&lt;span class="s2"&gt; passes&amp;#34;&lt;/span&gt;
&lt;/span&gt;&lt;/span&gt;&lt;span class="line"&gt;&lt;span class="cl"&gt; &lt;span class="nb"&gt;echo&lt;/span&gt; &lt;span class="s2"&gt;&amp;#34;&amp;#34;&lt;/span&gt;
&lt;/span&gt;&lt;/span&gt;&lt;span class="line"&gt;&lt;span class="cl"&gt; sed -ie &lt;span class="s2"&gt;&amp;#34;s/row_size=&lt;/span&gt;&lt;span class="nv"&gt;$prev_size&lt;/span&gt;&lt;span class="s2"&gt;/row_size=&lt;/span&gt;&lt;span class="nv"&gt;$SIZE&lt;/span&gt;&lt;span class="s2"&gt;/g&amp;#34;&lt;/span&gt; &lt;span class="nv"&gt;$EDIT_FILE&lt;/span&gt;
&lt;/span&gt;&lt;/span&gt;&lt;span class="line"&gt;&lt;span class="cl"&gt; sed -ie &lt;span class="s2"&gt;&amp;#34;s/col_size=&lt;/span&gt;&lt;span class="nv"&gt;$prev_size&lt;/span&gt;&lt;span class="s2"&gt;/col_size=&lt;/span&gt;&lt;span class="nv"&gt;$SIZE&lt;/span&gt;&lt;span class="s2"&gt;/g&amp;#34;&lt;/span&gt; &lt;span class="nv"&gt;$EDIT_FILE&lt;/span&gt;
&lt;/span&gt;&lt;/span&gt;&lt;span class="line"&gt;&lt;span class="cl"&gt; sed -ie &lt;span class="s2"&gt;&amp;#34;s/max_iterations=&lt;/span&gt;&lt;span class="nv"&gt;$prev_pass&lt;/span&gt;&lt;span class="s2"&gt;/max_iterations=&lt;/span&gt;&lt;span class="nv"&gt;$PASS&lt;/span&gt;&lt;span class="s2"&gt;/g&amp;#34;&lt;/span&gt; &lt;span class="nv"&gt;$EDIT_FILE&lt;/span&gt;
&lt;/span&gt;&lt;/span&gt;&lt;span class="line"&gt;&lt;span class="cl"&gt; sed -ie &lt;span class="s2"&gt;&amp;#34;s/repetitions=&lt;/span&gt;&lt;span class="nv"&gt;$prev_pass&lt;/span&gt;&lt;span class="s2"&gt;/repetitions=&lt;/span&gt;&lt;span class="nv"&gt;$PASS&lt;/span&gt;&lt;span class="s2"&gt;/g&amp;#34;&lt;/span&gt; &lt;span class="nv"&gt;$EDIT_FILE&lt;/span&gt;
&lt;/span&gt;&lt;/span&gt;&lt;span class="line"&gt;&lt;span class="cl"&gt;
&lt;/span&gt;&lt;/span&gt;&lt;span class="line"&gt;&lt;span class="cl"&gt; &lt;span class="nv"&gt;prev_size&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="nv"&gt;$SIZE&lt;/span&gt;
&lt;/span&gt;&lt;/span&gt;&lt;span class="line"&gt;&lt;span class="cl"&gt; &lt;span class="nv"&gt;prev_pass&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="nv"&gt;$PASS&lt;/span&gt;
&lt;/span&gt;&lt;/span&gt;&lt;span class="line"&gt;&lt;span class="cl"&gt; &lt;span class="nv"&gt;count&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="k"&gt;$((&lt;/span&gt;count+1&lt;span class="k"&gt;))&lt;/span&gt;
&lt;/span&gt;&lt;/span&gt;&lt;span class="line"&gt;&lt;span class="cl"&gt;
&lt;/span&gt;&lt;/span&gt;&lt;span class="line"&gt;&lt;span class="cl"&gt; ./&lt;span class="nv"&gt;$BUILD_FILE&lt;/span&gt;
&lt;/span&gt;&lt;/span&gt;&lt;span class="line"&gt;&lt;span class="cl"&gt;
&lt;/span&gt;&lt;/span&gt;&lt;span class="line"&gt;&lt;span class="cl"&gt; &lt;span class="nb"&gt;echo&lt;/span&gt; &lt;span class="s2"&gt;&amp;#34;&amp;#34;&lt;/span&gt;
&lt;/span&gt;&lt;/span&gt;&lt;span class="line"&gt;&lt;span class="cl"&gt; &lt;span class="nb"&gt;echo&lt;/span&gt; &lt;span class="s2"&gt;&amp;#34;Generated for &lt;/span&gt;&lt;span class="nv"&gt;$SIZE&lt;/span&gt;&lt;span class="s2"&gt; x &lt;/span&gt;&lt;span class="nv"&gt;$SIZE&lt;/span&gt;&lt;span class="s2"&gt; with &lt;/span&gt;&lt;span class="nv"&gt;$PASS&lt;/span&gt;&lt;span class="s2"&gt; passes&amp;#34;&lt;/span&gt;
&lt;/span&gt;&lt;/span&gt;&lt;span class="line"&gt;&lt;span class="cl"&gt; &lt;span class="nb"&gt;echo&lt;/span&gt; &lt;span class="s2"&gt;&amp;#34;&amp;#34;&lt;/span&gt;
&lt;/span&gt;&lt;/span&gt;&lt;span class="line"&gt;&lt;span class="cl"&gt;
&lt;/span&gt;&lt;/span&gt;&lt;span class="line"&gt;&lt;span class="cl"&gt; cp &lt;span class="s2"&gt;&amp;#34;./build/mandelbrot.png&amp;#34;&lt;/span&gt; &lt;span class="s2"&gt;&amp;#34;./images/&lt;/span&gt;&lt;span class="nv"&gt;$count&lt;/span&gt;&lt;span class="s2"&gt;.mandelbrot_&lt;/span&gt;&lt;span class="si"&gt;${&lt;/span&gt;&lt;span class="nv"&gt;SIZE&lt;/span&gt;&lt;span class="si"&gt;}&lt;/span&gt;&lt;span class="s2"&gt;_&lt;/span&gt;&lt;span class="si"&gt;${&lt;/span&gt;&lt;span class="nv"&gt;PASS&lt;/span&gt;&lt;span class="si"&gt;}&lt;/span&gt;&lt;span class="s2"&gt;.png&amp;#34;&lt;/span&gt;
&lt;/span&gt;&lt;/span&gt;&lt;span class="line"&gt;&lt;span class="cl"&gt; 
&lt;/span&gt;&lt;/span&gt;&lt;span class="line"&gt;&lt;span class="cl"&gt; &lt;span class="k"&gt;done&lt;/span&gt;
&lt;/span&gt;&lt;/span&gt;&lt;span class="line"&gt;&lt;span class="cl"&gt;&lt;span class="k"&gt;done&lt;/span&gt;
&lt;/span&gt;&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;&lt;/div&gt;&lt;figure class="doc-figure"&gt;&lt;img src="https://miro.medium.com/v2/resize:fit:1400/format:webp/1*Wu-I9mjifg7MeS8DIU_Q9g.png" alt="Generated images" loading="lazy" decoding="async"&gt;&lt;figcaption aria-hidden="true"&gt;Generated images&lt;/figcaption&gt;&lt;/figure&gt;&lt;p&gt;More iterations means better quality of the image, you can zoom in and see the small details and tiny fractal repetitions, but to see these clearly the resolution must be increased.&lt;/p&gt;
&lt;h2 id="the-winning-moment"&gt;The winning moment&lt;a class="heading-anchor" href="#the-winning-moment" aria-label="Link to this section"&gt;&lt;/a&gt;
&lt;/h2&gt;
&lt;p&gt;It was Day 2 and we both had absolutely nothing with us.
We sat at the back thinking what we would present the crowd with.
We had about 3 hours in our hands to build something.
While brainstorming, we were randomly playing with the Mandelbrot directory already given by Intel in their Virtual Private Server.
We ran the code with the proper command and was later told to rerun using &lt;code&gt;qbus -I&lt;/code&gt; which redirected us to another portal maybe through some kind of routing.
We then decided to create a bash script to randomise four values in the original program i.e. &lt;code&gt;row_size&lt;/code&gt;, &lt;code&gt;column_size&lt;/code&gt;, &lt;code&gt;max_iterations&lt;/code&gt; and &lt;code&gt;repetitions&lt;/code&gt;.
We were able to do this and now we had started generating images in a folder.
Some images with higher resolution or more repetitions took more time than the rest.
The longest time was taken by 16K resolution image with 1000 repetitions which took approximately 50 minutes (probably still wayy less than an average C++ or Python script generating the same iterations and resolution).&lt;/p&gt;
&lt;p&gt;We demoed our solution on a presentation made on Keynote.
We remarkably remember that we were the only team with the minimal interaction with Nandy Sir and also the only ones to receive a round of applause from him.
At the end of all prizes, 6 stood out of 13 and were congratulated on stage for the same.
Then came the cash prizes for the Top 3.
As soon as second was done, we thought the claps were consolatory of nature and we were already packing.
But then we realised, we were indeed 1st. (Out of excitement, &lt;strong&gt;I smacked the Krish’s Macbook Pro, like Davie504 would slap bass&lt;/strong&gt; 😂).
In the end we received 12K INR as a prize in the form of Amazon Vouchers.&lt;/p&gt;</content:encoded></item></channel></rss>