Second-year Information Management student · Taiwan · 繁體中文
I build small things, and every one of them started as a problem I actually had. I wanted Claude Code to watch a tutorial video for me and found that agents can't watch video at all, so I built something that lets mine do it. I couldn't work out which IELTS words I was weak on, so the study site I wrote surfaces them without being asked. I quit every expense tracker at the same step, so mine has no categories to pick.
All of it runs at $0. I'm a student with no budget for servers and no appetite for a monthly API bill, so if a free tier covers it I use the free tier, and if it can run locally it runs locally. That constraint has pushed me towards architectural choices I wouldn't have made with a credit card.
vid-for-agents — giving an agent a way to watch video
One command turns an Instagram Reel, a YouTube URL, or a file already on disk into the two things an agent can read: sampled frames and a transcript. yt-dlp fetches, ffmpeg samples frames at even intervals, and local whisper.cpp transcribes. Everything runs locally: nothing uploaded, nothing billed, nothing to revoke later.
The part I find genuinely interesting is the output format. The generated README embeds the full transcript inline, but only lists the frame filenames with their timestamps. Text is cheap for an agent to read; images are expensive. So the agent spends its own context deciding which frames are worth opening, rather than me deciding that for it in advance.
Shell · ffmpeg · yt-dlp · whisper.cpp
CountAgent — an expense tracker with a deliberately stupid capture side
Say "lunch 120" to Siri and the capture side does exactly one thing: it appends that line, verbatim, to a plain text file. No fields, no dropdowns, no login, no server to be down. It is too dumb to fail, which is the entire point. Later, at my computer, I say "sort it out", and an agent reads the backlog, works out the date, the type, the category and the payment method, checks the line isn't a duplicate, and appends a nine-column row to a CSV.
The rules for that path live in an 8 KB natural-language document rather than a parser, so changing one means editing a sentence of prose. Bulk invoice imports still run through a hand-written keyword table, because that input arrives already structured and has no ambiguity left for a model to resolve.
Python · Claude Code · iOS Shortcuts
CalAgent — the same split, moved to a calendar
CountAgent can afford to be lazy. An expense that sits in the inbox for three days produces the same ledger. A calendar entry cannot: if the meeting reaches the calendar after the meeting, the agent was asleep at the one moment it was supposed to help. So the capture side stayed dumb and the digest side stopped waiting for me. A launchd job watches the inbox file and wakes a headless Claude within seconds of a line landing. I say "明天下午三點跟教授 meeting" to Siri, and about fifteen seconds later it is in the right calendar, with an end time I never gave it.
The parsing was not the hard part. Relative dates resolve against the capture timestamp instead of the clock at digest time, because a line captured on Friday and digested on Sunday would otherwise move Saturday's shift to Monday. The unattended run also gets a strict allowlist: two scripts it may execute, one that only adds events and one that only reads, plus write access to data/ and the inbox and nothing else. I checked that by asking it to edit the add-only script and to run raw AppleScript, and both were refused.
Since the rules are prose, I test them by handing the repo to a model that has never seen it. The first run created two events, correctly refused three, and refused two it should have taken: a weekday named after that weekday has already passed, like "這禮拜五" said on a Saturday, and the rules never said which way to resolve it. Its report named four gaps like that, and all four were real.
Python · AppleScript · launchd · Claude Code
ielts-study-desk — an IELTS study site organised around what I get wrong
ielts-study-desk.pages.dev · nothing to sign up for; your data stays in your browser.
My problem with vocabulary was never memorising it. It was not knowing which words I was weak on. So the site isn't built around a word list. It's built around one behaviour: what you get wrong comes back on its own. Leitner five-box scheduling, with intervals of 1, 2, 4, 7 and 14 days. A right answer moves a card up one box; a wrong answer sends it straight back to box one, however far it had got.
The frontend is one 2,231-line HTML file with no framework and no bundler, and the deploy script is fifteen lines that copy two paths into dist/. The backend (Cloudflare Pages Functions and D1) is optional: the whole app works without an account. It exists because iOS Safari clears localStorage after seven days without a visit.
I didn't cut corners on the auth. PBKDF2 at 100,000 iterations with a per-user salt, constant-time comparison, and a dummy hash computed even when the account doesn't exist, so response timing can't reveal which usernames are real. Going back over this repo recently I found something worse than any of that: a user's own Gemini key was riding along inside the state blob that syncs to my database. It's now stripped before upload, and there's a test that fails if anyone ever puts it back. Five test files load the whole document into jsdom and run 104 assertions against it.
vanilla JS · Cloudflare Pages Functions · D1 · Gemini API
sean-rpg — turning "pay for your own life" into a daily dashboard
sean-rpg.pages.dev · installable PWA, works offline.
It tracks daily habits, but only one number on it matters: what share of my own living costs I currently cover.
Nobody else has to use this, which is exactly why I spent longer on the interface than on the features. A washi-paper ground, a vertical CJK masthead, and a vermilion seal pressed onto anything completed. Vermilion never means anything else in the system, and gold appears exactly once in the whole palette. I threw out my own first choice of ink colour after measuring it at 3.18:1 against the paper, which fails WCAG AA; the CSS comments still carry the numbers.
One deployment decision I'd defend in a review: Cloudflare Pages serves only public/ as the site root, so db/ stays in the repo and off the CDN. curl sean-rpg.pages.dev/db/schema.sql returns the index page rather than the SQL. The whole backend is one 75-line file writing to D1.
vanilla JS · Cloudflare Pages · D1 · Service Worker
Mostly write: JavaScript (vanilla; I rarely reach for a framework) · Python · Shell
Deploy on: Cloudflare Pages, Functions and D1, where the free tier is far more than one person needs
AI: Gemini API · local whisper.cpp · and Claude Code as both my development environment and a runtime my projects call into
A tool I don't use isn't finished. I still use all of these every week. The ones I stopped using aren't on this page.
Write less code where you can. CountAgent's capture side stores one line and stops, and every judgement about what that line meant is handed to the model. For anything a person types freehand, a hundred if-branches guessing at intent loses to a hundred rules written as prose. Where the input is already structured, the branches win, and CountAgent still uses them there.
seanlu2006@gmail.com · happy to talk about any of this.

