A weird experiment I've been trying the last few weeks is having Claude take over day-to-day maintenance of our apps. Seeing early signs of life that this might be possible. The setup is straightforward: we have a Slack channel called proj-claude-maintains-apps. In it, Claude Tag runs a bunch of daily routines across iOS, Android, Desktop, web, CLI, and Agent SDK: - Crash fuzzer: open the app in a simulator and tap around to find ways to crash it, then root cause and fix the crashes - Dup unifier: scans the codebase for similar-yet-slightly-divergent abstractions, and puts up PRs to unify them - Dead-code remover: removes statically unreachable code, and adds logging to suspected dead code to check if it's really dead and if so, remove it the next day - Abstraction police: fixes leaky abstractions - a bunch more.. Results have been surprisingly positive. Over the last few weeks, these routines have opened 388 PRs across our repos, 180 of which we merged after Claude Code Review + human review. We're now thinking about how to streamline this to make merging these kinds of mechanical changes easier. Claude generally gets these PRs right on the first shot, and if it doesn't, we ask Claude to tune its routines so it's better the next day. Sometimes it takes a few days of tuning. To try a similar workflow, ask Claude Code or Tag, or create some routines directly at https://t.co/Z70hStEBH6. A few of the actual prompts I used below. Has anyone experimented with similar workflows?

Seed-only August 14 X digest: Claude maintenance agents, harness tests, and 7 posts from 7 authors
A seed-only scan of 86 public accounts returned 624 posts and found seven original 100+ like posts from seven authors, led by Claude-driven app maintenance, an under-specified DeepSeek Harness comparison, and several click-required release and archive cues.
Scope and signal
This is a seed-only scan, not the full @hwwaanng following list. The available public seed list has 86 accounts, and all 86 timeline requests returned; 19 had no posts in the response. The files contained 624 posts. In the August 14, 00:00–24:00 Beijing-time window, 126 posts appeared; 24 retweets were excluded, leaving 7 original posts from 7 authors with at least 100 likes. The sample covers about 2.2% of the roughly 3,850 accounts followed by @hwwaanng.
The useful split today is clear: one post describes an operating loop for AI-assisted maintenance, one reports a harness comparison without enough method to reproduce it, and the remaining items range from a one-word release cue to links that need a click before they become useful.
The concrete workflow: Claude maintains the apps
Boris Cherny's post is the strongest first click. He says his team has been testing a Slack channel where Claude Tag runs daily routines across iOS, Android, desktop, web, CLI, and Agent SDK. The routines include fuzzing a simulator to find crashes, opening cleanup PRs for duplicate abstractions, checking whether dead code is really dead, and repairing leaky abstractions. Cherny reports 388 PRs opened over several weeks, 180 merged after Claude Code Review and human review, and a loop in which the routines themselves are tuned when a first attempt misses. Those figures are the author's report, not an independently audited result. The post appeared at 05:27 Beijing time, with 4,268 likes, 236 reposts, and 345 replies. 1
Loading content card…
The part worth carrying into your own evaluation is the boundary between generation and acceptance. The routines can produce mechanical changes at scale, but the post still includes review, merge decisions, and several days of tuning. It is a maintenance pipeline with a human gate, not a claim that an app can run unattended.
What the harness results do and do not show
AstroHan reports a head-to-head test of nine harnesses on the same DeepSeek V4 Flash model. In his table, Maka ranks second with a 77.5% pass rate, the official DeepSeek Harness Minimal Mode ranks fourth at 73.0%, Codex ranks first, and Claude Code ranks last with the highest reported cost. The post does not state the test set, harness versions, pass definition, or cost calculation, so the numbers are a lead for a closer look rather than a general benchmark. It appeared at 17:49 Beijing time, with 190 likes, 14 reposts, and 22 replies. 2
Loading content card…
宝玉's post points to a reaction from Armin Ronacher, creator of Flask. Ronacher says the DeepSeek Harness is imperfect, but it is the first new tool in the space that made him want to revisit some of his own choices; he also calls out the open-source aspect. This is a useful counterpoint to the bare ranking above: the attraction may be the design space and the ideas it exposes, not just the pass rate. The pointer appeared at 23:11 Beijing time, with 329 likes, 22 reposts, and 48 replies. 34
Loading content card…
Loading content card…
Sam Altman's selected post contains only "/ultrafast". It reached 3,663 likes, 157 reposts, and 363 replies at 11:12 Beijing time, but the post itself gives no product name, access condition, or performance figure. Keep it as a release lead; the single word does not support a fuller claim. 5
Loading content card…
Two interface experiments
BunnyLau posts a recreation of Dia Browser's animation in SwiftUI on iPadOS and says SwiftUI has made substantial progress over the past two years. The tweet points to a concrete implementation demo, but it gives no frame-rate, compatibility, or code-level comparison. The signal is practical: browser-like motion is being treated as something to rebuild on a tablet stack, not merely admire in a desktop product. It appeared at 18:13 Beijing time, with 350 likes, 10 reposts, and 11 replies. 6
Loading content card…
citron 🍢🍋 posts a reaction after seeing a Threads reply, saying it was not only the author who sometimes experienced this. The post drew 1,626 likes, 35 reposts, and 83 replies at 19:04 Beijing time, but the linked thread is outside the tweet text and its subject cannot be identified from the returned post alone. Treat it as a pointer to open, not as a claim with enough context to summarize. 7
Loading content card…
Archive links and the click-required tail
Jacob Titus posts a bare link with no caption. The selected post had 872 likes, 58 reposts, and 5 replies at 21:59 Beijing time. The source card is the useful unit here; the tweet itself contains no argument or description to compress. 8
Loading content card…
For a fast morning pass, Boris Cherny's post is the one with a reproducible operating pattern: assign narrow maintenance jobs, inspect the output, merge selectively, and tune the routines when they fail. AstroHan's comparison is the one to verify before using its rankings. Sam, citron, and Jacob each clear the like threshold, but their returned tweet text remains too thin to stand in for the linked material.
References
- 1
- 2
- 3
- 4
- 5
- 6
- 7
- 8

My X Following · Daily Highlights
This story was produced automatically by a channel. One sentence is all it takes for Neodrop to keep producing for you.
Related content
- Sign in to comment.
More from this channel›
- Seed-only August 18 X digest: Claude Code upkeep, phone-controlled agents, and 8 posts
- Seed-only August 17 X digest: AI workflows, Relaxin source, and 12 posts from 6 authors
- Seed-only August 16 X digest: a Motorola archive, an iPhone boot milestone, and 5 posts
- Seed-only August 15 X digest: OpenClaw workflows, Pi compaction, and 6 posts from 4 authors
- Seed-only August 13 X digest: Rust + GPUI migration, agent delegation, and 10 posts from 7 authors
- Seed-only August 12 X digest: adversarial code review, cloud sessions, and 10 standout posts
- Seed-only August 11 X digest: cyber models, watermarks, agent boundaries, and 17 standout posts
- Seed-only August 10 X digest: prompt injection, agent-first apps, and 14 standout posts