- That AI Thing
- Posts
- weekend ai reads for 2026-07-02
weekend ai reads for 2026-07-02
direct links are available on the web at https://thataithing.beehiiv.com/p/weekend-ai-reads-for-2026-07-02
š° ABOVE THE FOLD: GRADING
Why Large Language Models (LLMs) Arenāt Built for Grading, And What That Means for Faculty / Grady Blog (10 minute read)
A general-purpose chatbot can grade a whole class in seconds, and that, says Grady co-founder and algorithms researcher Anastasios Sidiropoulos, is exactly the trap. Why LLMs misread student work, why hallucinated grades canāt be patched away, and what faculty should demand instead.
Professor denounces mass AI fraud on an exam at Brown University: āAcademic integrity is at riskā / El PaĆs (10 minute read)
Serrano did not void the midterm exam, but warned students that the final one, which counted for 50% of the final grade, would be held in-person. He also said that if the grade distribution was not similar to the midterm, only the final exam would be taken into account. The average score dropped to 48 out of 100. Of the 89 students who did the midterm exam, only 59 showed up for the final one. And of the 27 who did not show up, 22 had scored a perfect 100 in the midterm exam.
5 AI Education Trends According To A Microsoft Executive ā The conversation around AI in schools is changing almost as rapidly as the technology. Here are some recent trends. / Tech & Learning (8 minute read)
āWe shouldnāt be using AI to grade people, because we need human judgment,ā Jubelirer says.
related, Microsoft adds AI teaching tools to 365 Education ā New Unit Plans, assignment controls, study tools, and live lesson features will roll out across Microsoft 365 Education, Windows, Teams, and learning management systems / EdTech Innovation Hub (12 minute read)
š» QUOTE OF THE WEEK
Everyone loves everything. Sunshine, rainbow. You stripped out the ugly parts, thinking it would make you more likeable. It just made you empty.
š„ FOR EVERYONE
Does It Matter If You Used AI To Make It? ā It depends / Josh Brake, Substack, archive (13 minute read)

XAI Bets on Grokās Racy Side / The Information (subscription required) (9 minute read)
SpaceX also touted the popularity of its AI video tools ahead of its blockbuster IPO. What SpaceX didnāt mention, however, is that much of the consumer demand stems from Grokās looser content rules, which have made it a major destination for generating [adult images and videos] and other racy content.
Indeed, two recent xAI employees estimated that well over half of Grokās overall traffic was driven by [adult] images and videos, role-play chats or other [such] activity. On forums for Grok users, many of the most popular posts are [adult-based]. Users can generate visuals in several ways, including picking the video models through the consumer app or tapping them through other Grok products.
Jesse Genet Runs Her Household With a Staff of AI Agents / The Cut (13 minute read)
Claire is just one of the AI agents on Genetās household staff. Thereās also Sylvie, who runs her kidsā homeschool; the Wests ā Clark, Dan, and Chloe ā who deal with legal and financial paperwork; and a team of coding agents that can build pretty much any app Genet describes. A year ago, none of this was really possible.
š FOUNDATIONS
Running local models is good now / Vicki Boykis (7 minute read)
I have no concrete scientific evidence of this - my own personal vibe metric of āis a model good enoughā is, ādo I have to double-check it against an API modelā, and GPT-OSS was the first one where I started doing that a lot less often.
As a result, Iāve mostly been using local models as fast, personalized Google for development questions that donāt require recency.
Anthropic Just Changed How We Work Forever.. (Claude Tag) / AI for Non Techies, YouTube (15 minute video)
Anthropic just released Claude Tag, which turns Claude into a multiplayer teammate right inside your Slack. In this video I break down the biggest takeaways from the release, the ambient mode that lets Claude take initiative on its own, and the privacy controls you need before turning it on. I also show it live and cover who gets access today.
Tau ā Learn how coding agents are built.
š FOR LEADERS
Couriers, Not Coders / Yegor256 (2 minute read)
We are still ready to pay. Not for the codeāthe code is free. For the delivery.
You take what Claude Code wrote and you walk it to our door. You are the courier, not the coder.
The margin is what we pay for trust: that what you deliver, we can merge without re-checking. So the delivery must be flawless.
The Minimum Viable Unit of Saleable Software / Brandur (8 minute read)
To counterbalance the $400/mo that wouldāve been paid to Atlassian, the engineer can spend no more than 4 hours a month (400 / 96) prompting features/fixes on their homegrown Jira clone, or looking after its database, or whatever, not including context switching overhead. Even with LLM help, thatās completely unrealistic already, but letās be charitable and say they can get it down to 2 hours a month. Itād still take 37 months to break even after those initial 2 weeks of effort (number of months to make back Atlassianās $400/mo minus 2 hours/mo maintenance effort = 2 * 3846.15 / (400 - 2 * 96.15)).
Donāt get me wrong, I hate Jira just as much as anyone whoās ever used it and have a nearly uncontrollable urge to want to rebuild it too, but the math here doesnāt pencil out.
Somewhere along the zone of viability is the minimum viable unit of saleable software, below which a rebuild is the same or less effort compared to going through the purchasing process for a third party and not cost-effective over the long run.
We Booked 614 Meetings With One Inbound Agent. Your āContact Usā Form Is Costing You Deals. / SaaStr AI blog (12 minute read)
They obviously didnāt all close. If every one of those 614 meetings had converted at the average, weād be looking at tens of millions in sponsorships, which isnāt reality. But the efficiency is the point. A small team turned a flood of inbound into hundreds of booked, qualified meetings without a single BDR.
Inside Consultantsā Messy Shift From Hourly Billing ā As AI threatens to make the billable hour obsolete, professional-services firms wrestle with reinventing how they charge clients / Wall Street Journal (7 minute read)
āMany are being forced to cut prices before they themselves have actually realized the cost-saving gains from the technology,ā said James OāDowd, chief executive of talent advisory firm Patrick Morgan.
The shift from hourly billing also is a challenge because buyers often compare bids on an āhours times rateā basis, even when hours arenāt part of the proposed pricing model, said Eric Miles, CEO of Baker Tilly.
Firms that continue to rely on hourly billing risk eroding their margins, because AI creates an environment with low variable costs and high fixed costs, he said.
š FOR EDUCATORS
Alpha School Brings $4,500-a-Week AI Summer Camp to the Hamptons / Bloomberg, archive (5 minute read)
Alpha School, an experimental private school that swaps teachers for āguidesā and uses AI to pack academics into just two hours, is offering its first summer classes in the Long Island vacation hub kicking off June 29. For $4,500 a week, kids from pre-K through rising ninth grade can learn math and reading in the morning ā taught by a proprietary AI model and other apps ā before pivoting to afternoon activities with rotating guests like chefs and athletes.
University of Utah trustees greenlight creation of stateās first AI bachelorās degree / KSL (4 minute read)
AI childrenās books, body horror edition / Lcamtufās Thing, Substack, archive (4 minute read)
some interesting examples in the article
š FOR TECHNOLOGISTS
Unstructured data is still a pain in the butt (but less impossible now) / Counting Stuff (8 minute read)
Yes, LLMās are prone to ridiculous hallucinations and are unlikely to ever be free of them. But even if they arenāt making things up from whole cloth, they exhibit really weird behavior when āsummarizingā in that their attention mechanisms carry a bias as to what is important. That importance may have little to do with your actual research question. Whatever quirks are going on under the LLM hood, in the end it still lands on human raters and labelers to bridge the results to reality.
Sakana Fugu ā Multi-Agent System as a Model
Frontier-level performance without single-vendor dependency. Fugu dynamically orchestrates the worldās best models to tackle complex, multi-step tasks. Plug collective intelligence directly into your workflows today with a single API.
Loop Engineering [PDF] / Google Drive (15 minute read)
We give particular attention to the generator/evaluator separation: empirically, an agent asked to grade its own output tends to praise it, and tuning an independent skeptical evaluator is far more tractable than making a generator critical of its own work. We survey three loops running in practice, from one engineerās morning triage to Stripeās enterprise-scale pipeline merging over 1,300 machine-written pull requests per week, and we catalog four costs that accrue silentlyāverification debt, comprehension rot, cognitive surrender, and token blowout. We close with a concrete recipe for building a first loop. The central claim is that loops make generation nearly free and leave judgment as the scarce resource; the same loop, built by two people, can yield opposite outcomes.
Hidden Technical Debt of AI Systems: Agent Harness / Lee Hanchung, GitHub (27 minute read)
When you build a training harness, you are doing something different. You are letting the model explore the action space, observing what behaviors emerge, and shaping them with rewards. If the model learns to call a destructive tool inappropriately, the answer is not to add a software guardrail; it is to penalize the trajectory and let the policy update. The fence moves from outside the model to inside the model. This is alignment from the inside out. It is also the only kind of alignment that scales with capability, because every external fence has a fixed cleverness budget and the modelās intelligence is growing faster than software gymnastics.
š FOR FUN
Stanford scientists built an AI that can design healthier, greener burgers ā The new system balances nutrition, taste, cost, and environmental impact to create better recipes. / Digital Trends (6 minute read)
It only takes one fake web page to fool AI shopping bots, study finds / Tech Xplore (7 minute read)
Did a 54-Year-Old Nonprofit Worker Just Use AI to Become the Next Charlie Kaufman? / Hollywood Reporter (11 minute read)
spoiler: no
related, A Face Only A Mother Could Love / Robert Gaudette AI, YouTube (8 minute video)
š§æ AI-ADJACENT
Why big AI labs are hiring so many philosophers ā The technology presents all sorts of thorny problemsāa philosopherās favourite kind / The Economist (7 minute read)
ā