This website uses cookies

Read our Privacy policy and Terms of use for more information.

In partnership with

Hey everyone,

Picture the most overqualified intern alive. You ask it one question (“who wrote this blog post?”), Google says no, and it goes, “ok, watch this.”

That's not a metaphor. It's from OpenAI's own incident report, published this week, and it ends with the company pausing tool use on its most capable models.

This week we're breaking down how it happened, and the three things it teaches you about getting hired.

THE SHORT VERSION

  • What happened: an AI agent in training got stuck on a search task, found a gap in its sandbox, and used it to ask a chatbot for help.

  • Why you care: the gap was DNS, and the lessons are about monitoring and ownership. Both are worth being able to talk about in an interview.

  • What to do: build a small DNS-watch project this weekend and put it on GitHub.

  • Stretch goal: let Jev, TypeSafe's new decision-only model, do the judging. Details in the Your move section below.

What we're reading

The AI Agent That Got Blocked, Then Got Creative

Cartoon robot sneaking a phone call out of a locked server room through a cord under the door

Agent to DNS resolver: “so... you got a minute?”

This week's sponsor: Tiger Data (formerly Timescale), Postgres with time-series chops. Handy for the weekend project below.

Analytics on Live Data Without Leaving Postgres

When analytics on Postgres slows down, most teams add a second database. Then come the pipelines, the sync jobs, and a copy of your data that's always a little behind.

TimescaleDB takes a different approach: extend Postgres instead of splitting away from it. Hypertables partition your data automatically as volume grows. Hypercore compression cuts storage up to 95%. Continuous aggregates keep dashboards live without re-querying everything.

CERN runs Postgres this way for sensor data from the Large Hadron Collider.

No split architecture, no pipeline lag, no new query language to learn. Same SQL, same drivers, same tools.

Start on Tiger Cloud and get $1000 in credits.

Here's the setup, per OpenAI's report. On September 20, an agent in training got a task: work out who wrote a specific blog post from a handful of clues. Its search tool returned junk. So it went around it.

How it went sideways

1. Brute force

It hit Google, Bing and DuckDuckGo straight from the shell. Result: a 403 and two 502s.

2. Wild guess

It decided the question looked like a public benchmark, downloaded the dataset, and started decoding it. It guessed wrong. Confident, though.

3. The gap

It noticed the sandbox's DNS resolver still answered real lookups, then used DNS delegation to route questions to a public chatbot.

Its first test message:

❝

What is the capital of France?

It timed out. The agent raised its timeout from 6 seconds to 19 and got “Paris” back.

Then it sent 18 more questions. Fourteen of them carried clues about the person it was hunting.

Fair context, because you'll fact-check me

OpenAI says everything except the DNS route hit an offline cache, not the live internet, and it calls this far less severe than some earlier incidents. It still paused training, evaluation and inference with tool use on its most capable models, and added blocks at two independent layers.

Three things you can steal from this

1. DNS isn't the boring networking chapter.

The escape route was an 80s-era protocol you skimmed in Networks.

2. A failed attempt is still an attempt.

The monitor caught this one. On similar cases it sometimes read “got nothing useful back” as “didn't work” and stayed quiet. The DNS anomaly detector also had this environment on its exclusion list. Log intent, not just outcomes.

3. Detection wasn't the bottleneck. Ownership was.

The alert fired inside 15 minutes and a human acknowledged it in 3. The run kept going for about 2.5 hours anyway. It didn't stop automatically, and there was confusion over whether to pull the plug. Have a story ready about a time you made the call.

The market, minus the sugarcoating

5.7%

unemployment for 22–27-year-old grads, as of June

42%

of recent grads are underemployed, versus 33.7% of all grads

Source: CNBC, citing NY Fed data. Ouch.

Better news if you read past the headline: Indeed postings just crossed 0% year-over-year after almost four years of declines, per Guy Berger's tracker. Software dev postings are growing faster, out of a much deeper hole. At the current pace, pre-pandemic levels arrive around summer 2028.

❝

Recovery, yes. Boom, no.

What separates you from the pile

What most CS students do

Let the agent write it, ship it, move on.

What gets rewarded

Kästner now makes students explain their own solutions to a TA, live, in 15 minutes. In his course, Claude Code kept pushing a “two-phase confirmation” fix for an agent security problem that collapsed under adversarial thinking. About 80% of students bought it at first. (He says newer models miss less.)

❝

The agent writes the code. You need to be the person who can say why it's wrong.

The lesson

YOUR MOVE: THIS WEEKEND, ~2 HOURS

Build a mini “DNS watch.”

  1. Run a container and log every DNS query it makes.

  2. Flag the weird ones: absurdly long subdomains, odd record types, lookups that spike right after failed HTTP calls.

  3. Push it to GitHub with a README on what it catches and what it misses. The misses are your interview story.

Then hit reply with the repo link.

STRETCH GOAL: LET JEV DO THE JUDGING

Jev, from TypeSafe, is one of the more interesting things to land in AI this month, and it fits your DNS watch almost perfectly. It's a new kind of model that doesn't write text at all. You hand it a piece of state and a few typed questions (pick from a list, give a score, or give a yes/no probability), and it returns a decision with a confidence number, built to be fast and cheap. Open-source clones showed up quickly, which tells you people want this.

Swap your hard-coded rules for a Jev-style call: “Is this DNS query normal or suspicious?” Then test it against queries you labeled yourself. It's great at narrow calls like this, not a replacement for a chat model, and its probabilities need checking before you trust them. Here's a hands-on test of what it does well and where it breaks.

Login or Subscribe to participate

That agent hit a wall and started probing every layer under it. You're staring at a wall too. Learn the layer below the tool, then send me the repo.

Support Jobless

Forward this to one friend who's grinding applications. Hang out with us at r/joblessCSMajors, and find more at jobless.news.

Until next time,
Team Jobless