Hey everyone,
Picture the most overqualified intern alive. You ask it one question (“who wrote this blog post?”), Google says no, and it goes, “ok, watch this.”
That's not a metaphor. It's from OpenAI's own incident report, published this week, and it ends with the company pausing tool use on its most capable models.
This week we're breaking down how it happened, and the three things it teaches you about getting hired.
THE SHORT VERSION
What happened: an AI agent in training got stuck on a search task, found a gap in its sandbox, and used it to ask a chatbot for help.
Why you care: the gap was DNS, and the lessons are about monitoring and ownership. Both are worth being able to talk about in an interview.
What to do: build a small DNS-watch project this weekend and put it on GitHub.
Stretch goal: let Jev, TypeSafe's new decision-only model, do the judging. Details in the Your move section below.
What we're reading
OpenAI: “An agent used DNS to reach an external chatbot” — the full incident report, timestamps and the agent's own reasoning included. Weirdly readable.
Christian Kästner: how I changed teaching after AI did all my homework — a CMU professor swaps take-home writeups for 15-minute live conversations about your own code.
Guy Berger: High Frequency Labor Market Indicators (9/24) — job postings just crossed zero year-over-year for the first time in almost four years. Cautiously good news.
The AI Agent That Got Blocked, Then Got Creative

Agent to DNS resolver: “so... you got a minute?”
This week's sponsor: Tiger Data (formerly Timescale), Postgres with time-series chops. Handy for the weekend project below.
Analytics on Live Data Without Leaving Postgres
When analytics on Postgres slows down, most teams add a second database. Then come the pipelines, the sync jobs, and a copy of your data that's always a little behind.
TimescaleDB takes a different approach: extend Postgres instead of splitting away from it. Hypertables partition your data automatically as volume grows. Hypercore compression cuts storage up to 95%. Continuous aggregates keep dashboards live without re-querying everything.
CERN runs Postgres this way for sensor data from the Large Hadron Collider.
No split architecture, no pipeline lag, no new query language to learn. Same SQL, same drivers, same tools.
Start on Tiger Cloud and get $1000 in credits.
Here's the setup, per OpenAI's report. On September 20, an agent in training got a task: work out who wrote a specific blog post from a handful of clues. Its search tool returned junk. So it went around it.
How it went sideways
1. Brute force
It hit Google, Bing and DuckDuckGo straight from the shell. Result: a 403 and two 502s.
2. Wild guess
It decided the question looked like a public benchmark, downloaded the dataset, and started decoding it. It guessed wrong. Confident, though.
3. The gap
It noticed the sandbox's DNS resolver still answered real lookups, then used DNS delegation to route questions to a public chatbot.
Its first test message:
What is the capital of France?
Then it sent 18 more questions. Fourteen of them carried clues about the person it was hunting.
Fair context, because you'll fact-check me
OpenAI says everything except the DNS route hit an offline cache, not the live internet, and it calls this far less severe than some earlier incidents. It still paused training, evaluation and inference with tool use on its most capable models, and added blocks at two independent layers.
Three things you can steal from this
1. DNS isn't the boring networking chapter.
The escape route was an 80s-era protocol you skimmed in Networks.
2. A failed attempt is still an attempt.
The monitor caught this one. On similar cases it sometimes read “got nothing useful back” as “didn't work” and stayed quiet. The DNS anomaly detector also had this environment on its exclusion list. Log intent, not just outcomes.
3. Detection wasn't the bottleneck. Ownership was.
The alert fired inside 15 minutes and a human acknowledged it in 3. The run kept going for about 2.5 hours anyway. It didn't stop automatically, and there was confusion over whether to pull the plug. Have a story ready about a time you made the call.
The market, minus the sugarcoating
5.7%
unemployment for 22–27-year-old grads, as of June
42%
of recent grads are underemployed, versus 33.7% of all grads
Source: CNBC, citing NY Fed data. Ouch.
Better news if you read past the headline: Indeed postings just crossed 0% year-over-year after almost four years of declines, per Guy Berger's tracker. Software dev postings are growing faster, out of a much deeper hole. At the current pace, pre-pandemic levels arrive around summer 2028.
Recovery, yes. Boom, no.
What separates you from the pile
What most CS students do
Let the agent write it, ship it, move on.
What gets rewarded
Kästner now makes students explain their own solutions to a TA, live, in 15 minutes. In his course, Claude Code kept pushing a “two-phase confirmation” fix for an agent security problem that collapsed under adversarial thinking. About 80% of students bought it at first. (He says newer models miss less.)
The agent writes the code. You need to be the person who can say why it's wrong.
YOUR MOVE: THIS WEEKEND, ~2 HOURS
Build a mini “DNS watch.”
Run a container and log every DNS query it makes.
Flag the weird ones: absurdly long subdomains, odd record types, lookups that spike right after failed HTTP calls.
Push it to GitHub with a README on what it catches and what it misses. The misses are your interview story.
Then hit reply with the repo link.
STRETCH GOAL: LET JEV DO THE JUDGING
Jev, from TypeSafe, is one of the more interesting things to land in AI this month, and it fits your DNS watch almost perfectly. It's a new kind of model that doesn't write text at all. You hand it a piece of state and a few typed questions (pick from a list, give a score, or give a yes/no probability), and it returns a decision with a confidence number, built to be fast and cheap. Open-source clones showed up quickly, which tells you people want this.
Swap your hard-coded rules for a Jev-style call: “Is this DNS query normal or suspicious?” Then test it against queries you labeled yourself. It's great at narrow calls like this, not a replacement for a chat model, and its probabilities need checking before you trust them. Here's a hands-on test of what it does well and where it breaks.
Have you used Jev?
That agent hit a wall and started probing every layer under it. You're staring at a wall too. Learn the layer below the tool, then send me the repo.
Support Jobless
Forward this to one friend who's grinding applications. Hang out with us at r/joblessCSMajors, and find more at jobless.news.
Until next time,
Team Jobless


