Glean 拾遗
← All issues
#018 Latest 9/21–9/27 Published Sep 28

Writing Faster Than We Can Read

This week's thread is simple: machines now write code faster than humans can read it. Anthropic's CI absorbed a 25x job increase in six months, pushing the bottleneck from writing code to PR review; a behavioral guide insists agents carry no responsibility and PRs should stay under 600 lines. Along the way, a $38B fund shows how it evaluates and governs frontier models, and a new release shifts the cost curve. But we are not here only to talk about speed. Bob Nystrom reminds us that overlong names are a symptom and that meaning cannot be outsourced; ACID and "make invalid states unrepresentable" both argue that the scarce resource was never output but constraints and structure. When generation becomes cheap, the crafts of design, review and governance get expensive. That is the sentence this issue points to.

12 picks 5 sections ~2 hr
Section 01

When the Pipeline Breaks First

3 / 12
claude.com · 10 min
01

How Anthropic scaled test impact analysis as agentic coding broke CIAgent 写代码把 CI 压垮:Anthropic 测试影响分析服务的三次续命与重构

Anthropic's CI absorbed a 25x increase in jobs over six months: Claude now writes 80% of the code, per-engineer quarterly output is 8x the 2021-2025 rate, and the test suite grew 10x while headcount barely moved. The bottleneck moved from writing code to PR review to CI, landing on the test impact analysis service that picks which tests run on each change. Because v0 needed a single writer to keep per-test history ordered, it ran as one process and could not be sharded. Three patches followed: doubling cores bought 70 days, per-package sharding 29 days, and daily restarts under a day. The rewrite moved history into an in-memory data store: any listener worker appends results to a journal and exits stateless, a small consumer rolls the journal into per-test history every few seconds, and the selector queries it. One engineer finished in three weeks.

www.piglei.com · 3 min
02

A Behavioral Guide to AI-Assisted ProgrammingAI 编程行为指南:Agent 不担责,PR 控制在 600 行内

A practitioner's behavioral guide to AI-assisted programming, deliberately tool-agnostic. Core claims: AI agents extend capability but bear no responsibility—humans remain the final owner of the code and should never commit what they cannot explain Feynman-style. Prefer collaboration over delegation; do not behave like a PM who only states requirements. Split PRs (under ~600 lines) and run an AI pre-review before opening one, since review cannot catch everything. Use plan mode to investigate before writing code. Prefer mature libraries over letting the model hand-roll utilities. Make changes verifiable and always verify. For junior engineers: quality over speed, participate in bug fixing instead of unsupervised delegation, design for a fixed 30 minutes before comparing with the AI's answer, read official docs, and shore up design patterns, DDD, security and concurrency.

skillselion.com · 10 min
03

i-have-adhd: a Claude Code output style that leads with the next actioni-have-adhd:让 Claude Code 先说动作、少铺垫的输出风格 skill

i-have-adhd by ayghri is an output-style skill that reshapes every Claude Code reply rather than adding knowledge: the first line is a command, path or snippet; multi-step tasks are numbered; state is restated each turn as "Step 3 of 5 done"; visible lists cap at five items; and the reply ends with one action doable in under two minutes. The whole payload is ten numbered rules plus a five-item pre-send deletion check that strips announcing openers, closing pleasantries, sidebars, empty hedges and idioms, shipped as v0.3.0 in a single SKILL.md with no tools and no MCP server. Install uses the two plugin commands from the repo INSTALL.md. Because the frontmatter sets disable-model-invocation: true, nothing changes until you type /i-have-adhd; staying on across sessions needs the SessionStart hook plus touch ~/.claude/.i-have-adhd-always. The catalog lists 14,235 installs, quality tier A and LOW risk, against a repo with 43,673 stars and 2,494 forks under MIT.

Section 02

The Hidden Debt of a Framework

2 / 12
www.piglei.com · 4 min
04

AI Coding Is a Framework, Not a LibraryAI 编程是框架还是库:抽象泄露与认知债务

Piglei argues that AI coding tools are better understood as a framework than a library. Frameworks own the program's overall structure and buy convenience at low upfront cognitive cost; AI tools do the same, with natural language replacing code as the input. Using Django REST Framework as the case study, he shows a four-line ModelViewSet generating a full CRUD API, then details what customizing a create response or adding list filters actually costs: rewriting get_queryset and stacking if/else patches. Dropping to a plain ViewSet makes the code longer but surfaces the hidden cognitive debt. Two framework problems persist with AI: abstraction leaks, when prompts fail and you must debug down to variable names, and loss of control, as in vibe coding where the agent owns the structure. His advice: treat AI as a library, find the prompt sweet spot, design the structure yourself, encode constraints in AGENTS.md, and review generated code.

www.piglei.com · 10 min
05

After 14 Years of Programming, Complexity Is Still the Enemy编程 14 年:好代码稀少,复杂度才是真正的敌人

After 14 years of shipping software, Piglei distills eight lessons on why programming stays hard. His central claims: good code remains a low-probability event even inside large companies serving millions of users; the ideal programming experience resembles solving LeetCode problems — separated concerns, fast precise feedback, near-zero cost of mistakes — and teams should chase that feeling through modularity, automated tests and shorter feedback loops; and the real enemy is not the product manager who keeps changing requirements but compounding complexity. Along the way he argues that SRP is really about people requesting changes, that perfectionism in code is a trap, that learning effort should be allocated by ROI rather than zeal, and that starting unit tests early reshaped his career. Written for engineers with production experience who are rethinking code quality and their own growth.

Section 03

Evals, Releases and Governance

2 / 12
www.anthropic.com · 23 min
06

Claude Opus 5.5: 40% cheaper, best alignment scoresClaude Opus 5.5 发布:成本降四成,对齐审计得分居首

Anthropic released Claude Opus 5.5, the first model in its new Claude 5.5 family, claiming Fable 5.1-level performance at 40% lower serving cost than Opus 5. Pricing drops to $4 per million input tokens and $20 per million output tokens, with cache reads at $0.20 (60% lower), and output generation is more than 30% faster. Anthropic reports its best automated behavioral audit scores to date across nearly 2,000 simulated scenarios, fewer attempts to cross containment boundaries, and stronger prompt-injection resistance. Because biology and cybersecurity capabilities approach Fable 5.1, the model ships with comparable safeguards: most cybersecurity tasks route to Opus 4.8, vetted labs can apply to a Life Sciences Verification Program, and preserved thinking blocks API context edits used for distillation. Benchmark tables carry standard errors and third-party caveats, and Anthropic concedes that score gaps no longer reliably predict real-world differences.

claude.com · 8 min
07

Balyasny: evaluating and governing frontier models at $38B scale380 亿美元基金的模型评测与 Agent 治理:Balyasny 访谈

An interview with Balyasny Asset Management's chief AI officer on how a $38B multi-strategy firm puts frontier models into production. BAM evaluates new models on thousands of real financial tasks with verifiable outcomes — equities, macro, commodities — both standalone and inside its own agentic environment using the same tools and files its users have, watching for numerical errors, missed coverage and retrieval failures. On the relevant subset Claude Fable 5 scored 89.4% versus 86.1% for the prior production model; a set of economics problems no model had ever solved finally passed, and BAM re-ran and independently re-checked the eval before accepting the result. Merger-arbitrage packages dropped from three-to-five days to under one, with a roughly 30-minute agent run and mandatory human review. Governance is framed as controls around the model — data boundaries, least privilege, tool-level permissions, logging, human approval — not as model selection. Note this is vendor-published customer material.

Section 04

Constraints Are Scarcer Than Output

4 / 12
kevinmahoney.co.uk · 6 min
08

Applying "Make Invalid States Unrepresentable" to Schema Design让非法状态无法表示:两个数据库建模案例

Two production cases of the "make invalid states unrepresentable" principle applied to schema design. First: modelling a contiguous timeline as a List (Date, Date) admits gaps and overlaps; storing only the split dates as a Set Date makes contiguity and non-overlap structural, and adding a split becomes a single set insert. Second: a contract system that stored both fixed and default contracts in one table allowed contract gaps — an optional end date plus a "default" flag made them easy to create, the per-contract mutation API guarded nothing, and real gaps reached production, costing hours to trace. Removing default contracts from the table and inferring them when no fixed contract exists eliminates both the gaps and the optional end date. The author traces the original design to object-oriented thinking that reifies every concept as a row rather than a proposition, and closes with two ways to forbid overlapping contracts: a database excludes constraint, or allowing overlaps in the write model and flattening them in a read-model projection terminated by the next contract's start date.

kevinmahoney.co.uk · 5 min
09

Consistency is Consistently Undervalued被长期低估的一致性:拆服务之前先算清 ACID 的账

The author argues that when teams adopt microservices or non-ACID databases, the most underrated loss is the transaction. Using a single constraint — one user must have exactly one profile — he shows how a single-database transaction gives atomicity and isolation for free: both rows are created together or not at all, and no other process ever observes a half-written state. Split users and profiles into two services and all of that disappears. He walks through the failure paths one by one: a failed profile write leaves an orphan user; the compensating delete can itself fail, or the process can die mid-way; even with 100% reliable services, concurrent threads can see the user without the profile and create a duplicate; foreign keys no longer hold, so profiles can point at deleted or changed user IDs; periodic cleanup of dangling rows still leaves a window of inconsistency. Application-layer compensation, he concludes, amounts to rebuilding a half-working, buggy distributed ACID database.

journal.stuffwithstuff.com · 20 min
10

What Color is Your Function?函数着色:async 如何把语言劈成两半

Using an invented language where every function is either red or blue, the author shows what function coloring really means: red functions are asynchronous. Red can only be called from red, calls need special syntax, and some core library functions are red-only. The root cause is async IO: you must unwind the entire C callstack back to the event loop, so callbacks, promises, async-await and generators all end up manually reifying the callstack on the heap via a CPS transform. Async-await fixes rule four but not the split: sync functions return values, async ones return Future/Task wrappers that still need awaiting. Go, Lua and Ruby avoid coloring entirely by suspending whole goroutines, coroutines or fibers instead of returning. C# escapes the same way by using threads. Java, notably, never had the problem.

journal.stuffwithstuff.com · 9 min
11

Long Names Are Long标识符过长也是病:四个删词规则砍掉名字里的水分

Working on Dart at Google, Bob Nystrom sits on the second layer of code review: the "readability" pass that checks style, idioms and documentation rather than behavior. What he sees there is identifiers that are far too long. Names can be too short, but the pendulum has swung — names over 60 characters dwarf the operations performed on them, force line breaks, and push authors toward nested expressions instead of locals. His four guidelines: drop words the type already implies (nameString → name, holidayDateList → holidays); drop words that don't disambiguate, tested by asking whether the name still means the same thing without them; drop words the surrounding class or method already supplies (AnnualHolidaySale.promote, not promoteHolidaySale); and drop filler like manager, data, state or engine. A deliberately awful waffle example is trimmed step by step from DeliciousBelgianWaffleObject down to Waffle.garnish.

Section 05

Meaning Cannot Be Outsourced

1 / 12
journal.stuffwithstuff.com · 16 min
12

The Value of Things: Utility, Meaning, and Generative AI效率的滑块:AI 造得出有用之物,造不出意义

Programming language designer Bob Nystrom separates value into utility — what a thing does — and meaning, which he argues comes from spending irreplaceable time in service of someone else and is not transferable. Generative AI sharply raises utility but cannot generate meaning: a ChatGPT-assisted screenplay written in a tenth of the time would carry a tenth of the meaning for his brother. He proposes a rough model in which efficiency is a slider setting the ratio of utility to meaning, and offers a heuristic — use AI freely for utilitarian work (he quotes a Washington Department of Ecology listing that endorses AI for boilerplate, test generation, and safe refactoring), but keep the machine out of work meant to carry human connection. He flags the model's limits: arbitrary hardship does not scale meaning, and the process itself can be joyful. Scope is deliberately limited to one person's individual use, setting aside AI's global effects.