weekend ai reads for 2026-07-02

šŸ“° ABOVE THE FOLD: GRADING

A general-purpose chatbot can grade a whole class in seconds, and that, says Grady co-founder and algorithms researcher Anastasios Sidiropoulos, is exactly the trap. Why LLMs misread student work, why hallucinated grades can’t be patched away, and what faculty should demand instead.

Serrano did not void the midterm exam, but warned students that the final one, which counted for 50% of the final grade, would be held in-person. He also said that if the grade distribution was not similar to the midterm, only the final exam would be taken into account. The average score dropped to 48 out of 100. Of the 89 students who did the midterm exam, only 59 showed up for the final one. And of the 27 who did not show up, 22 had scored a perfect 100 in the midterm exam.

5 AI Education Trends According To A Microsoft Executive — The conversation around AI in schools is changing almost as rapidly as the technology. Here are some recent trends. / Tech & Learning (8 minute read)

ā€œWe shouldn’t be using AI to grade people, because we need human judgment,ā€ Jubelirer says.

  • related, Microsoft adds AI teaching tools to 365 Education — New Unit Plans, assignment controls, study tools, and live lesson features will roll out across Microsoft 365 Education, Windows, Teams, and learning management systems / EdTech Innovation Hub (12 minute read)

 

šŸ“» QUOTE OF THE WEEK

Everyone loves everything. Sunshine, rainbow. You stripped out the ugly parts, thinking it would make you more likeable. It just made you empty.

ā€œUnstoryā€ (source)

 

šŸ‘„ FOR EVERYONE

Does It Matter If You Used AI To Make It? — It depends / Josh Brake, Substack, archive (13 minute read)

XAI Bets on Grok’s Racy Side / The Information (subscription required) (9 minute read)

SpaceX also touted the popularity of its AI video tools ahead of its blockbuster IPO. What SpaceX didn’t mention, however, is that much of the consumer demand stems from Grok’s looser content rules, which have made it a major destination for generating [adult images and videos] and other racy content.

Indeed, two recent xAI employees estimated that well over half of Grok’s overall traffic was driven by [adult] images and videos, role-play chats or other [such] activity. On forums for Grok users, many of the most popular posts are [adult-based]. Users can generate visuals in several ways, including picking the video models through the consumer app or tapping them through other Grok products.

Claire is just one of the AI agents on Genet’s household staff. There’s also Sylvie, who runs her kids’ homeschool; the Wests — Clark, Dan, and Chloe — who deal with legal and financial paperwork; and a team of coding agents that can build pretty much any app Genet describes. A year ago, none of this was really possible.

 

šŸ“š FOUNDATIONS

Running local models is good now / Vicki Boykis (7 minute read)

I have no concrete scientific evidence of this - my own personal vibe metric of ā€œis a model good enoughā€ is, ā€œdo I have to double-check it against an API modelā€, and GPT-OSS was the first one where I started doing that a lot less often.

As a result, I’ve mostly been using local models as fast, personalized Google for development questions that don’t require recency.

Anthropic Just Changed How We Work Forever.. (Claude Tag) / AI for Non Techies, YouTube (15 minute video)

Anthropic just released Claude Tag, which turns Claude into a multiplayer teammate right inside your Slack. In this video I break down the biggest takeaways from the release, the ambient mode that lets Claude take initiative on its own, and the privacy controls you need before turning it on. I also show it live and cover who gets access today.

Tau — Learn how coding agents are built.

 

šŸš€ FOR LEADERS

Couriers, Not Coders / Yegor256 (2 minute read)

We are still ready to pay. Not for the code—the code is free. For the delivery.

You take what Claude Code wrote and you walk it to our door. You are the courier, not the coder.

The margin is what we pay for trust: that what you deliver, we can merge without re-checking. So the delivery must be flawless.

To counterbalance the $400/mo that would’ve been paid to Atlassian, the engineer can spend no more than 4 hours a month (400 / 96) prompting features/fixes on their homegrown Jira clone, or looking after its database, or whatever, not including context switching overhead. Even with LLM help, that’s completely unrealistic already, but let’s be charitable and say they can get it down to 2 hours a month. It’d still take 37 months to break even after those initial 2 weeks of effort (number of months to make back Atlassian’s $400/mo minus 2 hours/mo maintenance effort = 2 * 3846.15 / (400 - 2 * 96.15)).

Don’t get me wrong, I hate Jira just as much as anyone who’s ever used it and have a nearly uncontrollable urge to want to rebuild it too, but the math here doesn’t pencil out.

Somewhere along the zone of viability is the minimum viable unit of saleable software, below which a rebuild is the same or less effort compared to going through the purchasing process for a third party and not cost-effective over the long run.

They obviously didn’t all close. If every one of those 614 meetings had converted at the average, we’d be looking at tens of millions in sponsorships, which isn’t reality. But the efficiency is the point. A small team turned a flood of inbound into hundreds of booked, qualified meetings without a single BDR.

Inside Consultants’ Messy Shift From Hourly Billing — As AI threatens to make the billable hour obsolete, professional-services firms wrestle with reinventing how they charge clients / Wall Street Journal (7 minute read)

ā€œMany are being forced to cut prices before they themselves have actually realized the cost-saving gains from the technology,ā€ said James O’Dowd, chief executive of talent advisory firm Patrick Morgan.

The shift from hourly billing also is a challenge because buyers often compare bids on an ā€œhours times rateā€ basis, even when hours aren’t part of the proposed pricing model, said Eric Miles, CEO of Baker Tilly.

Firms that continue to rely on hourly billing risk eroding their margins, because AI creates an environment with low variable costs and high fixed costs, he said.

 

šŸŽ“ FOR EDUCATORS

Alpha School, an experimental private school that swaps teachers for ā€œguidesā€ and uses AI to pack academics into just two hours, is offering its first summer classes in the Long Island vacation hub kicking off June 29. For $4,500 a week, kids from pre-K through rising ninth grade can learn math and reading in the morning — taught by a proprietary AI model and other apps — before pivoting to afternoon activities with rotating guests like chefs and athletes.

AI children’s books, body horror edition / Lcamtuf’s Thing, Substack, archive (4 minute read)

  • some interesting examples in the article

 

šŸ“Š FOR TECHNOLOGISTS

Yes, LLM’s are prone to ridiculous hallucinations and are unlikely to ever be free of them. But even if they aren’t making things up from whole cloth, they exhibit really weird behavior when ā€œsummarizingā€ in that their attention mechanisms carry a bias as to what is important. That importance may have little to do with your actual research question. Whatever quirks are going on under the LLM hood, in the end it still lands on human raters and labelers to bridge the results to reality.

Sakana Fugu — Multi-Agent System as a Model

Frontier-level performance without single-vendor dependency. Fugu dynamically orchestrates the world’s best models to tackle complex, multi-step tasks. Plug collective intelligence directly into your workflows today with a single API.

Loop Engineering [PDF] / Google Drive (15 minute read)

We give particular attention to the generator/evaluator separation: empirically, an agent asked to grade its own output tends to praise it, and tuning an independent skeptical evaluator is far more tractable than making a generator critical of its own work. We survey three loops running in practice, from one engineer’s morning triage to Stripe’s enterprise-scale pipeline merging over 1,300 machine-written pull requests per week, and we catalog four costs that accrue silently—verification debt, comprehension rot, cognitive surrender, and token blowout. We close with a concrete recipe for building a first loop. The central claim is that loops make generation nearly free and leave judgment as the scarce resource; the same loop, built by two people, can yield opposite outcomes.

Hidden Technical Debt of AI Systems: Agent Harness / Lee Hanchung, GitHub (27 minute read)

When you build a training harness, you are doing something different. You are letting the model explore the action space, observing what behaviors emerge, and shaping them with rewards. If the model learns to call a destructive tool inappropriately, the answer is not to add a software guardrail; it is to penalize the trajectory and let the policy update. The fence moves from outside the model to inside the model. This is alignment from the inside out. It is also the only kind of alignment that scales with capability, because every external fence has a fixed cleverness budget and the model’s intelligence is growing faster than software gymnastics.

 

šŸŽ‰ FOR FUN

Stanford scientists built an AI that can design healthier, greener burgers — The new system balances nutrition, taste, cost, and environmental impact to create better recipes. / Digital Trends (6 minute read)

 

🧿 AI-ADJACENT

Why big AI labs are hiring so many philosophers — The technology presents all sorts of thorny problems—a philosopher’s favourite kind / The Economist (7 minute read)

 

ā‹„