Every day, The Neuron teaches readers one practical AI skill they can use immediately. This second September collection includes every published AI Skill of the Day from September 16 through September 30, 2026.
Skim the headings, grab the prompts that fit your work, and return when you need a better way to use Claude, ChatGPT, Gemini, Codex, or an AI agent.
How to use this digest
- Skimming? Each entry starts with the practical outcome.
- Trying one? Copy the included prompt or workflow and replace the bracketed details.
- Catching up? The skills are ordered by their original newsletter date.
🎓 September 16
Add confidence thresholds to agent decisions
One underrated idea in JEV is not the speed. It's the probability attached to the answer.
You can use the same pattern in your own AI workflows today: make the model return a decision plus a confidence score, then route low-confidence cases to a human or a stronger model.
Try this structure:
- Ask for a fixed decision, not an essay.
- Require a confidence score from 0 to 1.
- Set an escalation rule, e.g. anything below 0.85 gets reviewed.
Prompt:
Classify this request as APPROVE, REVIEW, or REJECT. Return only the decision, a confidence score from 0 to 1, and one sentence explaining the most important uncertainty. If confidence is below 0.85, choose REVIEW.
It will not magically calibrate a general-purpose model the way Jev’s RLCD process is designed to, but it gives your workflow a useful escape hatch instead of pretending every answer deserves the same level of trust.
🎓 September 17
Turn every edit into a reusable rule
In the above story, the agent was taking notes you didn’t ask it to (and certainly didn’t want it to). Now flip that scenario: say you do actually want the AI to keep one thing, the feedback you already gave it. How do you do that?
Every's Katie Parrott calls this Compound Writing: each correction you give AI should improve the next draft.
After you edit a draft, for example, you have the model compare its version with yours and extract only lessons that should be applied again. Save those rules in one instruction file and reuse it every time. Alongside the core concept, Every also published the open plugin behind the workflow.
Sample Prompt version:
Compare your draft with my edited version. Extract only reusable rules that would improve future drafts. Organize them under Voice, Structure, and Content. Ignore one-off factual corrections. Write each rule as a short instruction I can reuse.
🎓 September 18
Stress-test a prompt before users do
Before shipping an agent prompt, simulate the ugly cases: vague users, conflicting requests, missing data, and long conversations. Respan Prompt Simulations generates realistic users and scenarios, then runs full multi-turn conversations against a committed prompt so you can see exactly where it breaks.
- Commit the prompt version.
- Generate realistic edge-case users and scenarios.
- Review failures, revise, and rerun.
🎓 September 20
Use Jev when the answer is a choice, not an essay
Most AI calls do not need a chatbot. TypeSafe’s Jev is built for more bounded decisions: a simple yes/no, a score, a category, or which choice to pick from a list.
This weekend popped off with devs going wild with Jev included demos that ranged from sorting 500 emails to screening thousands of listings. Jev handles tiny decisions; ChatGPT or Claude can handle the messy exceptions.
Other good tests:
- Romàn scored 700 sales leads and personalized outreach in about 40 seconds for $0.09.
- A Postgres demo used Jev inside a WHERE clause to judge 129 rows in about one second for $0.0009.
- Browser experiments showed Jev making decisions for fractions of a cent, including finding a flight in roughly seven seconds for $0.0039.
The pattern across these demos is useful: Jev works best when you already know the shape of the answer and need to make the same small judgment hundreds or thousands of times. Instead of asking a frontier model to reason through every row, email, lead, or button, use the cheap decision model for the routine cases and save the expensive model for ambiguity.
These are developer demos, not standardized benchmarks, but they point to a handy rule: if the answer is basically yes, no, this one, that one, or give it a score, try a small decision model before reaching for the biggest chatbot.
Here’s how you can use it in your own systems:
- Define the allowed answers before the model runs.
- Batch lots of small decisions together.
- Send low-confidence or open-ended cases to a frontier model.
Check out more demos and POVs / tips re: Jev in today’s Around the Horn Digest!
🎓 September 21
Make Your AI Automation Safe to Retry
Your AI agent updates a CRM, sends an email, or submits an order. Then the connection times out. The action may have succeeded even though the agent never got confirmation.
The fix lives inside your automation, immediately around the step that takes the real-world action:
AI decides → check if already done → perform action → record success
In a tool like n8n, that means:
- Before the Gmail, CRM, payment, or HTTP action, create a stable ID from something that won’t change, like the lead ID or order number.
- Check that ID against a Data Table, database, or the destination itself. If it already exists, stop.
- If it doesn’t, run the action and save the ID as completed. Any retry checks the same ID before acting again.
If the service supports idempotency keys, you can pass that stable ID directly with the request. n8n’s guide shows how to do this with its HTTP Request node, retry controls, Data Tables, and error handling. See n8n’s full retry-safe workflow guide
There’s also a copyable n8n workflow template that puts the check before payments, emails, database writes, or other actions.
Rule to steal: before an automation repeats an action, make it prove the first attempt didn’t already work.
🎓 September 22
Don’t /compact your agent just because you’re taking a break
Long coding-agent sessions can feel messy, so /compact looks like housekeeping. Don’t treat it that way.
Kun Chen points out that manual compaction usually triggers a summarization pass over the whole current context. On a huge session, that can itself be expensive, and the compacted session may have to rebuild context without the same cache savings.
Better rule: let the agent harness auto-compact at its tuned threshold. Codex, for example, automatically compacts once its token limit is exceeded. Use manual compaction only when context pressure is actually hurting the task, not before lunch because the context meter looks ugly.
🎓 September 23
Build a two-tier model stack
Do not make your most expensive model do every part of an agent job. Split the work by decision quality.
- Use your strongest model to write the technical plan, architecture, and acceptance criteria.
- Hand well-scoped implementation tasks to a cheaper model, and parallelize where the tasks are independent.
- Bring the result back to the stronger model for code review, security review, or final synthesis.
Copy/paste:
Plan this task in phases. Prioritize quality, but do not be wasteful. Identify which steps need frontier-level judgment and which can be delegated to cheaper subagents. Write clear acceptance criteria for every delegated step, then review the combined result for correctness, security, and missed requirements.
🎓 September 24
Fan out research, then funnel it back down
One chat is good at following one line of thought. However, big research questions often have too many independent branches for that. Borrow the pattern from Anthropic's enzyme search: split the work into parallel scouts, then collapse the pile into a short list.
- Break your research question into independent lanes, such as competitors, evidence for, evidence against, pricing, user reports, or technical constraints.
- Give every scout the same output shape: claim, evidence, caveat, source link, and confidence.
- Send every scout result to one reviewer that removes duplicates, challenges weak evidence, flags conflicts, and ranks the few findings worth your attention.
You are basically replacing "one giant research prompt" with a tiny newsroom.
Copy/paste:
You are coordinating a research swarm. Break this question into 5 independent research lanes. For each lane, return only: claim, evidence, caveat, source link, confidence. Then merge the lanes, remove duplicates, flag conflicts, and rank the 5 findings that most change the answer. Do not hide disagreements or weak evidence.
🎓 September 25
Clean out conflicting AI instructions
You know how we tell you to save useful corrections so AI stops making the same mistake? That advice comes with a housekeeping problem. Every's Katie Parrott had saved old outlines, writing feedback, and several versions of her essay templates so her assistant could remember what worked. Eventually, the drafts started coming back crowded and flat. When she looked through the files, she found that alternative templates had turned into simultaneous requirements for every piece. The assistant was trying to follow all of them.
So Katie ran a review across the instructions, archived the old folders, and rebuilt two current guides: one for the column's structure and one for her voice. She also stopped saving every intermediate version. The idea is to give the assistant a smaller set of instructions that actually agree with each other.
If you have a writing project that has gotten worse after months of tweaking, try the same cleanup. Ask AI to identify conflicting or outdated rules with exact file references, decide what still applies, then archive the rest. Run a familiar assignment afterward and compare the draft. Sometimes your AI needs a closet cleanout more than another pep talk.
Copy/paste:
Review the instructions and examples in [project/folder]. Identify duplicated, conflicting, and outdated rules, citing the specific files and short excerpts. Propose what to keep, archive, or rewrite. Do not modify files until I approve.
FROM OUR PARTNERS
🎓 September 27
Run your product team like a research lab
Dan Shipper's advice for surviving nonstop model upgrades is basically: stop making the same people explore the frontier and execute the roadmap. Those are opposite jobs. Exploration means trying lots of weird stuff and throwing most of it away. Product work means focus, reliability, and saying no.
His setup at Every is tiny. One or two people can be the lab. The useful pairing is a "pirate" who rapidly builds messy experiments to find value, plus an "architect" who steps in once something starts working and turns it into a real system.
The important part is how ideas graduate:
- Try multiple approaches in parallel and expect roughly 90% to die.
- Dogfood the survivors on real work. Ask: is this actually useful, or just new?
- Only harden what people keep using, then test whether it's dramatically better and affordable enough to scale.
Every's copy-editing experiment, "KateBench," is the concrete version. Once their editor actually started using it, they built a dashboard around accepted suggestions and the work she still had to do afterward. Shipper said it cut that remaining editing work by 12% month over month.
That's the filter I like: don't promote the demo because it looks futuristic. Promote it because, a month later, people still want it.
🎓 September 28
Benchmark the harness, not only the model
ARC Prize just gave us a clean example of why model comparisons can mislead. Gemini 3.8 Flash scored 10.37% on ARC-AGI-3 with a standard harness, then 35.0% with a provider adapter around the same model and reasoning level.
A harness is the software around the model that manages memory, tool calls, context, and what gets carried from one step to the next.
The better setup preserved Gemini's hidden reasoning state and compacted context instead of repeatedly starting from a flatter view of the task. Same brain, better workspace. WindTunnel found a similar pattern with browser agents: giving them WebMCP tools changed speed, cost, and success.
Turns out "which model?" can be the wrong first question.
- Pick 10 real tasks you actually care about, not a generic benchmark.
- Freeze the model, reasoning level, and task instructions. Change one harness variable, such as memory, context compaction, or the tool interface.
- Score completed tasks, human rescues, total cost, and elapsed time. The useful winner is the setup that finishes more real work per dollar.
Copy/paste:
Help me compare two harnesses for the same AI model. Use these 10 tasks: [tasks]. Keep the model, reasoning level, and task instructions fixed. For each run, record success, retries, human rescues, total tokens/cost, elapsed time, and failure mode. Then tell me which harness improved completed work per dollar, not which one looked smarter.
🎓 September 29
Branch a good ChatGPT thread instead of starting over
Ever get 20 messages deep and want another direction without wrecking the thread?
On ChatGPT web, you can “branch” from any message. The new chat keeps everything before it; the original stays untouched.
- Hover over the message.
- Click More actions (⋯), then Branch in new chat.
- Try the alternate plan, rewrite, or debugging path.
Basically Git branches if you’re technical, but for the conversation you were afraid to touch. OpenAI says it works for logged-in web users, including Projects. Try it!
🎓 September 30
Make the AI prove it understood you first
Lauren Tan shared one of her most-used prompts, and it's useful because a lot of bad AI work starts before the model writes a single word: it misunderstood the assignment.
A model can execute the wrong interpretation perfectly. Asking it to restate your goal surfaces that mismatch before you burn time, tokens, or 14 tool calls solving the wrong problem.
Try this before a complicated research, coding, planning, or writing task:
restate in your own words what you think my goals are and what the problem I'm trying to solve is
If the restatement is wrong, correct it before the model starts. If it's right, you've just given yourself a cheap comprehension check before the expensive work begins.
Very small prompt. Very high chance of preventing a very dumb afternoon.
Keep learning
That closes out September Part 2 and the month's AI Skill of the Day collection.
Have a workflow you want us to unpack next? Request an AI Skill of the Day.