Bombe Bench
Every Enigma simulator lets you encrypt. Almost none let you do what Bletchley Park actually did, which was break it. Bombe Bench starts with the machine: three of the five rotors, a reflector, ring settings, up to thirteen plugboard leads, and the double step of the middle rotor that the textbooks get wrong as often as not. Type a letter and the current is drawn through the plugboard, the three rotors, the reflector and back, with the full wiring faint behind it, and a slider replays any keypress. The cycle structure of each rotor and of the whole stack is printed underneath, because the fact that the machine is always thirteen swapped pairs (so no letter ever becomes itself) is the crack the attack uses. Then you switch to the Attack tab. Paste an intercept, type a crib like WETTERVORHERSAGE, and slide it: every offset where a letter would encrypt to itself goes red. The surviving offset becomes a menu, a little graph with one edge per crib position, and the loops in it light up in colour. Press Run and a Web Worker tries all 60 rotor orders and 17,576 positions, following each loop around the composed scramblers (Turing’s test) and then flooding the implications through Welchman’s diagonal board, with a running count of positions rejected. A million positions take about a second, and what comes out is a short table of stops with the plugs each one implies and a checking-machine verdict. A tracer lets you pick any position and watch one hypothesis spread until the test register lights twice, with the diagonal board switchable off so you can see why loops mattered so much before Welchman. Instructors generate challenges from a bank of stereotyped WWII traffic and hand out a link; the page grades the decryption letter by letter.
The idea asked for React and Tailwind; this is plain TypeScript and SVG, and it builds in half a second. The honest limitation is the one the real Bombe had: it assumes the middle rotor does not step during the crib, and it reports positions with the ring settings at A. The challenge generator guarantees the first condition, but a “hard” challenge with random rings decrypts correctly through the crib and then drifts after the first turnover, and sorting out the ring setting is left to the student with no help beyond a hint. The checking machine is also just a consistency check on the implied plugs, not a model of the historical one. 55 vitest tests, 98.8% statement coverage on the core.
See README for more details on the code or try it at https://bombe-bench.netlify.app
(Last update October 2026)
State Space Atlas
Search algorithms get taught on toy grids, and students never see the shape of the space being searched. State Space Atlas draws the whole thing. Pick a puzzle (Towers of Hanoi, the 8-puzzle and its smaller cousins, the classic Klotski, a couple of Rush Hour boards, or your own JSON definition) and a Web Worker enumerates every reachable state: 181,440 for the 8-puzzle, 13,011 for Klotski once identical pieces and mirror images are merged. Each state becomes a dot in the column of its distance from the start, so the silhouette of the drawing is the frontier size at every depth and “exponential growth” is something you can point at. Then press Play. BFS floods the columns one at a time, DFS threads a thin line, iterative deepening re-sweeps the shallow states round after round, greedy darts, and A* carves a narrow channel toward the goal, with live counts of expanded states, frontier size and peak memory, and a table that keeps every run on the same atlas side by side. The part I like most is the heuristic editor. You type the body of a JavaScript function, and because the atlas already knows the true distance of every state, it checks admissibility and consistency on all of them, reports the worst overestimate, and paints the offending states red. The presets include one heuristic that is inadmissible on purpose and one for Hanoi that is admissible but not consistent, which is a distinction most students only meet as a sentence in a textbook. Instructors can export a puzzle and an expansion target as a link and ask students to beat it.
The idea called for WebGL and a graph library; this is a single 2D canvas with edges cached in an offscreen bitmap, and it draws the 8-puzzle in a few milliseconds. The honest limitation is the cap of 400,000 states, which keeps the 15-puzzle and bigger block boards out of reach, and the layout only has two shapes, layered and radial, so a dot’s horizontal position tells you its depth and nothing else. Heuristics also run on the main thread, so a slow one on the 8-puzzle freezes the page for a couple of seconds while it checks. 65 vitest tests, 99.6% statement coverage on the core.
See README for more details on the code or try it at https://state-space-atlas.netlify.app
(Last update October 2026)
Reverse Minesweeper Studio
Minesweeper’s worst moment is the forced guess: you have done everything right and the last two cells are a coin flip. Reverse Minesweeper Studio turns the game around. You paint a picture with mines (a heart, your initials, a logo), choose the safe cell a player opens first, and a solver plays the board from that opening using only logic. It climbs the same ladder a good player does: single numbers first, then two overlapping numbers compared against each other, then every possible mine layout around a region, and finally the total number of mines left, which is the only thing that clears the inside of a walled box. If it finishes, the board is certified guess-free and gets a difficulty rating from the hardest step it needed. If it gets stuck, the cells where a player would have to gamble turn amber, two mine layouts that both fit every visible number can be flipped back and forth, and a background search tries toggling each nearby cell and lists the smallest edits that fix it. A slider replays the deduction one step at a time, which makes it a decent first lesson in constraint satisfaction. A finished board travels in the link itself, and the friend who opens it plays to uncover your drawing.
The idea suggested React and Tailwind; this is plain TypeScript and a single canvas so a 40 by 40 board stays smooth while you drag. The honest limitation is that the fix search only tries one-cell edits near the trouble. Some boards need two coordinated changes, and then the list offers the edit that gets furthest and you fix it one step at a time. The difficulty meter is also coarse: it reports the hardest kind of reasoning needed, not how long a person will take, so a board with one pairwise step rates the same as one with fifty. 40 vitest tests, 99.9% statement coverage on the core.
See README for more details on the code or try it at https://reverse-minesweeper-studio.netlify.app
(Last update October 2026)
Elevator Algorithm Lab
Operating systems courses teach disk scheduling as lists of cylinder numbers: FCFS, SSTF, SCAN, LOOK, and the students dutifully add up head movement without ever feeling why starvation matters. Elevator Algorithm Lab makes the same algorithms physical. A policy is one JavaScript function, dispatch(car, state), that sees the car, the other cars and the people waiting on each floor, and returns the next floor to go to. Run it and an animated building plays the trace: cars climb, doors open, orange dots pile up on floors nobody is serving. Each run reports mean, 95th percentile and maximum wait (the last one is the starvation number), energy as floors travelled, stops, and how many passengers were still waiting when time ran out. Pick a second policy and both replay in sync on the same traffic with the better number highlighted. Then flip to the disk view and the identical run is redrawn as cylinder across, time down, one zigzag per car: the diagram from the textbook, with the lesson that SCAN on a building is SCAN on a disk made literal. Six presets come as readable starting code, and on the textbook trace they move exactly 640, 236, 236 and 208 cylinders. Instructors write scenarios as small JSON files with seeded traffic generators (morning rush, lunch spike, a trap where one person on floor 1 waits while a stream rides between 30 and 40 and SSTF never comes down), can mark a scenario hidden for grading, and share everything by link.
The idea called for Supabase to hold scenarios and a class leaderboard. There is no server: the leaderboard lives in the browser, stores only names, scores and an eight-character hash of the policy, and travels as CSV or JSON that students hand in or merge. The honest limitation is that “hidden” is hidden in the interface only; the scenario is compressed inside the share link, so a student who decodes the URL can read the traffic. The model is also deliberately simple: a car goes straight to the floor the policy named and does not stop for anyone on the way, which is exactly what a disk head does and what makes the two views the same algorithm, but it is less like a real elevator than a controller that picks up hall calls en route. 53 vitest tests, 100% statement coverage on the core.
See README for more details on the code or try it at https://elevator-algorithm-lab.netlify.app
(Last update October 2026)
Where Does It Say
Students ask the same logistics questions all term (“can I use AI on lab 3?”, “is the late penalty per day?”), and a general chatbot answers them confidently even when the syllabus says nothing. Where Does It Say turns that around: it may only answer by pointing at a sentence. The instructor drops in the syllabus, assignment specs and course policies (PDF, Markdown or plain text), and every answer is a passage copied from them, checked word for word against the document and highlighted in place in a side-by-side viewer, with the section and page it came from. If nothing qualifies, it says “Not stated in the course documents,” names the words in the question the documents never use, and logs the question. A student can also mark a real quote as not answering the question. Students export their logs, and the instructor loads them into a gap report: unanswered questions grouped by week and by similar wording, plus the words students keep using that the syllabus never contains. A second tab checks quotes pasted from anywhere, such as a chatbot’s answer; a changed quote (“10% per week” for “10% per day”) is shown next to the closest real passage with the altered words marked. The whole course pack travels compressed inside the share link, so there is no server and nothing is uploaded.
The idea called for Claude Haiku behind a serverless function to propose quotes, with code verifying them. I kept the verifier and replaced the model with keyword retrieval (BM25 over sentences, a small synonym table for course logistics, and matching that knows “lab 3” falls inside “labs 1 to 4” but not “labs 5 to 8”). It is free, instant and gives the same answer every time. The limitation is that a verified quote is real but not necessarily relevant, and a question worded very differently from the syllabus can be refused even when the answer is there. Both cases end up in the gap report, which is the point. PDF support extracts the text layer, so scanned syllabi need OCR first. 91 vitest tests, 99.7% statement coverage on the core.
See README for more details on the code or try it at https://where-does-it-say.netlify.app
(Last update September 2026)
Cipher Ladder
A frontier model cracking a 370-year-old cipher made the news this month, which is a good excuse to teach the old craft. Cipher Ladder is a six-rung course in classical cryptanalysis: Caesar, affine, Vigenere, columnar transposition, monoalphabetic substitution, and homophonic substitution, each unlocked only after the previous one is solved. Beside the ciphertext sits the workbench a student actually needs: letter, bigram and trigram frequencies drawn against the English baseline, index of coincidence for every candidate key length with the English and random lines marked, a Kasiski table of repeated trigrams and the periods that divide their distances, a per-column chart for Vigenere, and keyboard-only substitution and symbol boards that update the decryption as you type. Hints come from a budget the instructor sets and each one reveals a single piece of the key. After a solve, a deterministic solver attacks the same ciphertext and its step log appears next to your own path: which shift chi-squared preferred and whether the “tallest bar is E” guess held, where Kasiski and IC disagreed, every read order tried for the columns, how the trigram hill-climb converged, and, for the homophonic rung, exactly which letters it still had wrong. Instructors build their own ladder (cipher, key, text, hint budget per rung) and the whole assignment is encoded in the link; students download a progress file and a dashboard reads any number of them.
The idea called for an LLM debrief and a Supabase roster. I replaced the LLM with a written-out program because a program can show its evidence, is free, and fails honestly: on the homophonic rung it usually leaves a couple of rare letters swapped and says so. The limitation is the language model behind the solvers, which is trained on only about eleven thousand letters of the bundled public-domain passages. It is calibrated for texts of a few hundred letters; give the substitution or homophonic solver a short instructor text and it will overfit garbage into fluent-looking English, which the confidence score now discounts but does not cure. There is no server, so the dashboard only knows about the files a student hands in. 65 vitest tests, 100% statement coverage on the core.
See README for more details on the code or try it at https://cipher-ladder.netlify.app
(Last update September 2026)
Repo Weight Map
Every term a student repo arrives with a virtualenv, a node_modules, a 400 MB dataset, or a build folder committed, and nobody notices until a push fails or a clone takes ten minutes. Repo Weight Map is the diagnostic: pick a folder in the browser and it draws a zoomable treemap of where the bytes are, colored by what they are (source, dependencies, generated output, data, media, the .git directory itself), then runs nineteen rules that know what junk looks like in Python, Node, JVM, .NET, Unity, C++, Rust, and Terraform projects, including a committed venv spotted by its pyvenv.cfg, bin/ and obj/ only when a .csproj sits next to them, and Unity’s Library/ only under a folder that also has Assets/ and ProjectSettings/. Each finding explains why the files do not belong in git, shows the bytes at stake, and contributes its .gitignore lines. Untick the ones that are wrong for your project and the before/after estimate and the diff against your existing .gitignore update. A standalone HTML report can be downloaded and attached to a help request. Files never leave the tab: only names and sizes are read, plus the root .gitignore so it does not suggest lines you already have.
The treemap is a hand-written squarified layout rather than d3, and there is no backend, no API, and no runtime dependency. The limitation is the half of the idea I skipped: history mode, which would read the .git packfiles to find large blobs that were deleted from the tree but still live in every clone, and print the filter-repo command to purge them. That needs a delta-resolving pack reader in the browser and it did not fit the weekend. The app does count .git and will tell you, usually correctly, that it is bigger than the cleaned tree, but it cannot tell you what is inside. 40 vitest tests, 100% statement coverage on the core.
See README for more details on the code or try it at https://repo-weight-map.netlify.app
(Last update September 2026)
CSS Scheduler
Every year our teaching coordinator places about 80 sections across roughly 50 instructors, and the working document is a spreadsheet. The constraints that actually bite are mundane: nobody can teach two courses in the same slot, someone is on sabbatical in Winter, a teaching professor owes a set number of courses. CSS Scheduler does not try to make the decision. It makes the constraints visible while a human makes it. Instructors sign in with their UW Google account and submit their own preferences (which quarters, how many sections, each course rated eager, willing, reluctant or unqualified, preferred days, times and modality, and how many new preps they will take on). A pure conflict engine then checks a draft schedule against twelve rules, from hard errors like a double-booked instructor or room down to info notes like an unwanted modality. It does no I/O, so it can rerun on every drag.
The seed data came from parsing three real quarterly time schedules (Autumn 2025 through Spring 2026): 240 sections, 199 of them staffed. That turned up two things guesswork would have got wrong. CSS runs on a fixed grid of six two-hour blocks against MW and TTh, with single days for hybrid sections, and 21 people taught CSS courses without appearing on the faculty page, so the instructor table is editable rather than scraped from the directory. Teaching load is a rank baseline (8 for teaching track, 5 for tenure track) minus release rows, each with a reason and optionally pinned to a year, so a chair’s standing release and a one-year grant buyout both still make sense a year later. There is no backend server: React on Netlify, Supabase for Postgres, Google OAuth and row level security (19 tables, 37 policies). An RLS regression test and a 20-check end-to-end test sign in a throwaway account through the real auth API and try to escalate its own role. Sign-ins are logged by a trigger on the auth table rather than by the app, so a client bug cannot skip them.
The honest status: two of five planned iterations are done. Preference collection works; the drag-and-drop assignment board the conflict engine was written for comes next, then reporting. There is no solver, by choice, though the data model leaves room for one, and reminder emails are not built. Login is restricted to uw.edu accounts, so it is not open to the public. The repo also keeps a verbatim log of the first build session: every prompt, every question asked back, and the options I turned down, which is often the more useful half of a decision record.
See README for more details on the code or try it at https://uwb-css-scheduler.netlify.app (UW login required).
(Last update September 2026)
Agent Blast Radius
Every MCP server, browser AI sidebar, and coding-agent allow rule was a reasonable yes on the day you clicked it. Added up, most of us have no idea what our AI tools can now reach. Agent Blast Radius reads the configs you already have (Claude Desktop, Cursor, VS Code and Claude Code MCP files, Claude Code permission settings, Chrome extension manifests, VS Code extension package.jsons) and scores every tool from 0 to 100 across five dimensions: filesystem, shell, network, credentials, and browser data. A filesystem server pointed at your home directory is not the same as one pointed at a project folder. A Docker server with the Docker socket mounted is root on your machine. A Bash(curl:*) allow rule is broader than it looks. Each capability comes with one plain sentence on what could go wrong and the exact bit of config that triggered it. Along the way it catches plaintext API keys, database passwords sitting in command-line arguments, unpinned npx packages that fetch whatever was published last night, and tools set to auto-approve. You drop the files into the page, nothing gets uploaded, and you can download the report as a single HTML file.
The idea asked for a local CLI that finds the files by itself. I built a static web page instead, so there is nothing to install, but the price is that you have to go find the files yourself (the page lists where each tool keeps them). The honest limitation is that it only scores what a config declares. It never reads a server’s code, so it recognizes known servers by name and treats anything it does not recognize as a program running with your full privileges. That is the right default, but a scary fork with a friendly name will look friendly. 95 vitest tests, over 99% coverage, no backend, no API keys.
See README for more details on the code or try it at https://agent-blast-radius.netlify.app
(Last update September 2026)
YPDSA
A tutor that hands a data structures student working code is worse than no tutor at all. YPDSA is the digital teaching assistant for my CSS 342/343 courses, and the interesting part is what it refuses to do. Answer-withholding is enforced as a per-turn, machine-checkable contract rather than a line in a system prompt: a non-LLM policy core reads trusted learner state and sets a ceiling on an eight-rung help ladder, a deterministic detector strips solution code out of replies, and a separate judge checks each risky reply against the contract before it is sent. The design, and the method used to tune it against scripted student personas, are written up in Teaching an LLM Tutor to Withhold the Answer.
What it does say comes from my own course rather than the internet. An offline pipeline ingests the real materials (existing PowerPoint lectures, Whisper transcripts of my own recordings, the textbook, assignment specs, Canvas PDFs, the course GitHub repos, even the course Discord) into roughly 12,000 indexed documents across nine collections, extracts a voice profile from them, and generated all 40 lecture decks plus a 40-chapter textbook the app serves next to the chat. Progress is gated by mastery: 136 learning goals, each with a pool of eight questions, where failing to show a prerequisite blocks a code attempt but never a conceptual question, and a deferred exam snoozes for a day instead of following the student around. Every reply carries a thumbs up and down, a thumbs down asks what was wrong, and anything flagged reaches me by email the next morning with a deep link straight to that message in the transcript.
Two honest notes. This is the one project I run on a real Anthropic API key rather than the claude CLI, so every student turn costs money and clearing that key is the kill switch. The infrastructure is otherwise deliberately cheap: Render plus a free-tier Supabase Postgres that pauses itself after a week of no traffic, which took the login page down once before a keep-alive job and a database health indicator went in. The pilot numbers are modest and stated as such, 100 of 100 prompts answered without errors and 31 of them citing a specific slide, with correctness spot-checked on a sample rather than verified across the whole run. Login is whitelisted, so this is available to UW students at this point and not to the public.
Source is not public for this one. Try it at https://ypdsa.pisan.me/ (login required).
(Last update September 2026)
MyChessMaster
Stockfish already knows which of your moves were bad. What it cannot tell you is why you played them, and that is the part a club player actually needs. MyChessMaster splits the job. The engine evaluates every position (MultiPV 3) and flags the moves that lost real win probability; the coach model is handed only those moments, with the FEN, the move played, the engine’s top lines, the phase and the clock, and is told to reason from those lines and invent nothing. It returns a reusable pattern name, an error category, one paragraph of explanation, and the question you should have asked before moving. Even time pressure is stamped from the %clk comments rather than asked of the model. That division is the whole design: the engine owns the truth, the model owns the language.
Everything then accumulates. Mistakes become spaced-repetition drills built from your own games, and the protocol matches the error: conversion and endgame failures are played out against a strength-limited engine until you hold the win, calculation errors demand the line several moves deep, opening deviations become flashcards, and quiet decoy positions from your own games are mixed in so detection gets trained and not just solving. Correct means within 3 win-probability points of best, strict in a balanced position and forgiving in a decided one. Import an opponent’s games (FIDE export, lichess or chess.com username) and you get a dossier cut to the colour they will have against you, a prep sheet where every claim cites its evidence, punish drills from their mistakes, and a Prepare page that plans the deck by days to go: the whole deck far out, the misses and top lines at four days, the sheet quiz on the day. Afterwards a debrief reports which prepared positions came up and which critical moments fell outside every line you had prepared.
It all runs on your machine. Games stay in a local data/ folder and the only thing that leaves is the text of a critical moment, sent to the claude CLI on a normal subscription, so there is no API key and no per-game cost. The limits follow from that: analysis depth is your hardware’s problem, the public site is a read-only mirror published from the machine doing the analysing, and this is a club of three plus an admin, not a product. Scouting needs an opponent with public games; with none it falls back to a prep built from your own games against players in that rating band. And the whole thing is tuned for one player at FIDE ~2000, so the thresholds, drill protocols and tone of the explanations all assume someone who knows why a move is bad once you point at it.
Source is not public for this one. Try it at https://mychessmaster.net/
(Last update September 2026)
Prompt Shrink Ray
Every production prompt accumulates cruft: “please,” “in order to,” the same instruction restated in three sections because nobody remembers writing it the first time. Prompt Shrink Ray takes the whole thing apart into labeled sections (system, context, examples, task) by explicit label or by shape (a You are... opener, matched Input:/Output: pairs), then compresses each one at a level you pick: strip filler words, simplify verbose phrasing (in order to becomes to), drop a sentence that already appeared earlier in the prompt, or trim a long run of few-shot examples down to two with a note of how many got cut. Fenced code is hidden behind a placeholder before any of that runs, so a rule never touches the code you actually pasted. What comes back is a live token-and-cost estimate, a word-level diff of exactly what moved, and a retention score that names every number, quoted string, or proper noun the compression dropped, so “it looks shorter” comes with a receipt.
The idea wanted two Claude calls: one to compress, one to judge whether the compressed version still meant the same thing. Both are rule tables instead, so the same prompt always compresses the same way, offline and free, and “did it still mean the same thing” becomes a narrower but checkable question: did it keep every number and code fragment, not a fuzzy judgment call from a second model. Token counts use the characters-over-4 rule of thumb providers themselves quote for English text, not a real tokenizer, so they are for comparing before and after, not for an actual invoice. 76 vitest tests, 100% statement coverage, no backend, no API keys.
See README for more details on the code or try it at https://prompt-shrink-ray.netlify.app
(Last update September 2026)
Summit Navigator
Conference programs are laid out for desktop monitors and read in hotel hallways on phones. Summit Navigator takes the inaugural ACM AI Leadership Summit (Hyatt Regency Atlanta, Aug 30 to Sep 2, 2026) and turns the program into something you can drive with a thumb: day tabs pinned to the top, sessions grouped by start time, eleven color-coded track chips, search that matches titles, speakers, rooms, and descriptions, and a star on every session that builds a “my agenda” view saved in the browser. A now/next banner works off the conference clock in Atlanta no matter where you are, and one toggle flips every listed time into your own timezone with exact IANA math. The whole conference lives in one JSON file the app validates loudly on startup, so pointing it at a different conference is a data edit, not a rewrite.
The honest limitation is the data itself: the build machine could not reach aisummit.acm.org (network egress policy), so the bundled schedule is reconstructed from ACM’s public announcements. Real venue, real dates, real tracks and keynote speakers (LeCun, Barto, Brooks), illustrative session times and rooms, and the footer admits as much. Also, the summit ended the day before this was built. Oh well: the app is the reusable part. 42 vitest tests, 100% coverage on the core logic, no backend, no API keys.
See README for more details on the code or try it at https://summit-navigator.netlify.app
(Last update September 2026)
Cron Cartographer
15 4 * * 1-5 tells you nothing at a glance. Paste it here and it comes back as “At 04:15 on Monday through Friday” plus a calendar heatmap of the next 30, 90, or 365 days with every firing marked; hover a day for the exact times. It reads five-field cron (names, steps, wrap-around ranges like FRI-MON, the @daily shortcuts), a working subset of RRULE, a pasted GitHub Actions schedule: snippet, or plain English: type “every weekday at 4:15am” and get the cron expression back, with a note for every assumption it made. Two timezone pickers, one for where the job runs and one for where you are, and the daylight-saving edges are handled the way cron actually behaves: the 2:30am job in Los Angeles simply does not fire on March 8, and the page tells you so, while the ambiguous 1:30am in November fires once, not twice. Even the odd corner where 0 0 13 * 5 means the 13th OR any Friday is honored. A share link carries the whole thing in the URL.
The idea wanted the Claude API for the English-to-cron part and npm packages for the parsing. Hand-written rules do both: same phrase, same cron, offline, free, zero runtime dependencies, and the timezone math builds exact offset tables from the browser’s own IANA data instead of trusting a library. The limitation is the usual one for rules: the English parser knows a fixed set of phrasings, and anything past them gets “could not understand” rather than a guess. Ask for “every 2 weeks” and it tells you plain cron cannot say that, which no cron string ever warned anyone about. 74 vitest tests, 92% coverage, no backend, no API keys.
See README for more details on the code or try it at https://cron-cartographer.netlify.app
(Last update September 2026)
Type Witness
Paste a TypeScript snippet and the compiler’s reasoning plays back as a story, one narrated step per expression. const x = 42 keeps the literal type 42 while let y = 42 widens to number, and the step says why. A call to a generic function shows the declared signature, the binding the compiler picked (T = string), and the resolved signature side by side. The unannotated word in ["a"].map(word => word.length) explains that its type flowed in from the array, and inside if (typeof v === "string") the same v that was declared string | number shows up narrowed, with the declared and narrowed types both on the card. Compiler errors are threaded into the story right after the step where inference went astray, so you can rewind to the exact moment the code and the compiler stopped agreeing. Hover any expression for its type, click it to jump to its step, or hit play and watch the whole thing unfold. Students see the type system think; I have watched plenty of them treat it as a wall that says no.
For once the idea had nothing for me to replace: no LLM, no backend, just the real TypeScript compiler (pinned at 5.6.3) running in a Web Worker with all 57 ES2022 lib files bundled in, so analysis is exact, offline, and nothing you paste leaves the browser. Every sentence of narration is filled in from what the checker actually computed, so the explanation cannot drift from the code. Two honest caveats: recovering the T = string bindings reads internal compiler state that the public API does not expose (it degrades gracefully if that ever changes shape), and shipping a compiler to the browser costs about 1 MB gzipped on first load. 62 vitest tests, 97.6% coverage, no API keys.
See README for more details on the code or try it at https://type-witness.netlify.app
(Last update August 2026)
Prompt Genome
Paste a prompt and it comes back as a strand of typed genes: role, persona, context, task, constraint, format, example, each color-coded and each showing why it got its label (“forbids something”, “imperative verb”). A lint pass scores the genome 0 to 100 and points at the weak spots: no task, “various things, etc.”, asking for brief and comprehensive in the same breath, constraints that only say what to avoid. Every gene offers three canned rewrites (Harden it, Soften it, Make it checkable), previewed with the changed words highlighted, and a side-by-side diff shows the prompt as pasted against the prompt as edited. Good genes go into a library with tags and search, and a share link carries the whole genome inside the URL hash, so there is no server to keep alive.
The idea wanted Claude to do the parsing and write the alternatives. I used rules instead: cue patterns for segmentation, hand-written templates for the mutations, an LCS diff for the comparison. Same paste, same genome, offline, free. The limitation follows directly: the classifier only knows the cues it ships with, so an unusual sentence lands in “context” with an honest “no strong signal” note, and the mutations are mechanical rephrasings, not fresh ideas. The idea also wanted to diff model responses, and without API calls there are no responses to diff; the prompt diff and the lint score stand in for that. 87 vitest tests, 99% coverage, no backend, no API keys.
See README for more details on the code or try it at https://prompt-genome.netlify.app
(Last update August 2026)
Code Analogy Forge
Paste a code snippet or just type “recursion”, pick your audience (curious child, high school student, CS undergrad, non-technical adult), and get three analogies from three different everyday domains, each with a small table mapping code terms to the analogy. Switching the audience swaps the entire text, not just the vocabulary: the child hears about a present that is a box inside a box inside a box, the undergrad gets the same imagery tied to base cases, unwinding, and stack frames. A detector works out what your paste shows from keyword mentions and code shapes (lo/hi/mid bounds for binary search, .push() plus .pop() for a stack), and for recursion it extracts the actual function body, by brace matching or Python indentation, to check the function really calls itself. Save the good ones to a library with tags and search, copy any card as Markdown for slides, or send a read-only share link that needs no server: the link encodes ids into the URL hash and resolves against the corpus shipped with the app.
The idea called for Claude to write the analogies and Supabase to store them. I wrote the analogies up front instead: 47 concepts times 3 domains times 4 audiences is 564 short texts, every mapping checked by a person, same input same output, offline, free. localStorage replaces Supabase. The limitation is the flip side of the corpus: the forge knows exactly 47 concepts (variables through heaps, HTTP, garbage collection, and operating systems), and anything outside them gets an honest “no known concept detected” rather than a fresh analogy. The original pitch wanted the open-ended tail, and that part really did need the LLM. 125 vitest tests, 99.9% coverage, no backend, no API keys.
See README for more details on the code or try it at https://code-analogy-forge.netlify.app
(Last update August 2026)
HAR Detective
Export a HAR file from the DevTools Network panel (right-click, “Save all as HAR”), drop it on the page, and get the analysis you normally do by squinting at a wall of requests: an interactive waterfall plus a ranked list of what is actually wrong. Ten detectors cover the classics: N+1 loops (/api/products/1 through /8 becomes one finding with the batch-endpoint fix attached), static assets served without cache headers, hundreds of KB of JSON that never met gzip, redirect chains, API calls running single-file when they could run together, duplicate fetches, failing requests, slow server think-time, oversized payloads, and HTTP/1.1 origins paying DNS and TLS over and over. Each finding shows the guilty requests, the wasted bytes or milliseconds, and a copy-paste fix; click one and the rows light up in the waterfall. The whole report exports as Markdown for a PR.
The idea called for Claude to read the request data and write the report. HTTP performance problems are a small, well-known catalogue, so hand-written rules do the job: same HAR, same report, offline, and the file never leaves the browser (HAR exports are stuffed with cookies and session tokens, so this matters more than usual). The limitation is that rules judge shape, not intent: a detector can see that three API calls ran strictly one after another, but not that your app needed the first response to build the second request, so some findings are the app working as designed. It flags; you decide. 56 vitest tests, 100% coverage, no backend, no API keys.
See README for more details on the code or try it at https://har-detective.netlify.app
(Last update August 2026)
Claude for STEM Professors
A guide for faculty who have seen the demos and want to ship something: from zero to a deployed course app, no programming background required. The bet behind it is that the hard part for professors is not prompting, it is plumbing. So the first half is accounts, tokens, and connectors (GitHub, Netlify, Render, Canvas), and only then come three destination projects with paste-ready starter prompts: a course website live on Netlify in under an hour, an auto-graded practice-problem app students use before exams, and a Canvas assistant on Render that drafts announcements from live roster and assignment data.
Two things I insisted on. The guide is dated, not evergreen: everything was verified against vendor documentation in August 2026, the version sits at the top, and when a button wanders, the linked docs win. And token hygiene is a full section rather than a footnote, written from experience: tokens are passwords, a Canvas token carries FERPA-protected student data, and every paste-a-token session ends with rotate or delete. Every starter prompt closes with “Ask me any questions before you start”, the guide’s single highest-value habit, because Claude’s guesses are plausible and wrong in exactly the ways that waste your afternoon. Ships as one README plus a printable PDF (pandoc and weasyprint) with the URLs spelled out for reading away from a screen.
See README to get started; the printable PDF is in the repo.
(Last update August 2026)
SQL Replay
Type a SELECT, paste some CSV (or use the bundled customers/orders tables), and watch the query execute one stage at a time: FROM, JOIN, WHERE, GROUP BY, HAVING, SELECT, DISTINCT, ORDER BY, LIMIT. Every stage shows the actual rows: joined rows merge with their partners, LEFT JOIN survivors get NULL padding, rows that fail WHERE are struck through with the evaluated condition sitting next to them, groups collapse with their member lists. A narrator explains each stage from the real execution counts (“3 rows pass, while 2 rows fail and are removed”). Students pick up SQL syntax in a week and then spend a quarter with no mental model of what the database does with it. Watching the rows move is the model. The whole workspace, query plus data, is encoded into the URL hash, so a puzzle query goes to a class as a plain link.
The idea suggested an optional Claude API narrator. I made it the mandatory non-Claude narrator instead: every sentence is a template filled with numbers the executor just computed, so the explanation cannot drift from the screen, and the semantics underneath are the real thing (three-valued logic, Kleene AND/OR, COUNT(x) skipping NULLs, inclusive BETWEEN). The limitation is the point: it replays the logical clause order on tables capped at 200 rows, with no planner, no indexes, and no subqueries. A real engine would almost never execute a query this literally, so this teaches what a query means, not how a database makes it fast. 86 vitest tests, 99.7% coverage, no backend, no API keys.
See README for more details on the code or try it at https://sql-replay.netlify.app
(Last update August 2026)
Sound Sketchpad
Type “muffled explosion heard from underground” and it plays one. The description runs through a word-to-DSP recipe book: 20 base sounds (explosion, coin, laser, rain, bell, heartbeat) and 18 modifiers (muffled, tiny, metallic, underwater, retro), each modifier rewriting the recipe by re-filtering, re-pitching, or bolting on ringing partials. The app shows exactly which words matched, draws the waveform, lists every layer’s signal chain (oscillator to filter to envelope to gain), and gives you sliders for pitch, length, brightness, echo, and volume. One button exports a WAV at 16 or 24 bit. Same words, same sound, every time; a Variation button re-rolls only the noise seed.
The idea called for Claude to turn the description into Web Audio code. I replaced it with the recipe book, and skipped Web Audio for synthesis too: a pure sample-by-sample DSP engine (swept oscillators, seeded noise, ADSR envelopes, biquad filters, echo, bit crush) renders the exact buffer the browser plays, which means the tests assert on the actual samples you hear and the WAV export is a byte-for-byte re-render. The limitation is the vocabulary: 20 bases and 18 modifiers cover the game-jam staples, but describe anything outside them and you get a generic whoosh and a hint about words that work. The open-ended tail really did need the LLM. 59 vitest tests, 100% coverage, no backend, no API keys.
See README for more details on the code or try it at https://sound-sketchpad.netlify.app
(Last update August 2026)
Code Review Gauntlet
A code snippet appears with one to three planted defects and a countdown. Click the lines you would flag in review, submit, and everything is revealed: the off-by-one you caught, the SQL injection you sailed past, each with its root cause and fix. Scoring punishes spraying: 100 points per defect found, minus 25 for every clean line flagged, and the time bonus only pays out when you found them all. A radar chart tracks accuracy across logic, null-safety, security, performance, and style, so after a dozen rounds you can see which category you keep missing. There is also a daily challenge: the date is hashed into the puzzle seed, so everyone gets the same round with no server behind it.
The idea called for Claude to generate endless fresh puzzles and Supabase for a leaderboard. I skipped both. Puzzles come from a seeded mutation engine: 11 hand-written clean snippets (JavaScript, Python, SQL) and a catalogue of 51 single-line defects, from inverted guards to timing-unsafe hash comparisons, each rewriting exactly one line so the generated code always parses. Same seed, same puzzle, offline, free. The honest limitation is the flip side: 51 defects is not an infinite supply, and a regular player will start recognizing templates within a week or two. The original pitch really did need an LLM for that part. 46 vitest tests, 99.8% coverage, no backend, no API keys.
See README for more details on the code or play it at https://code-review-gauntlet.netlify.app
(Last update August 2026)
UI Diff Lens
Drop in before and after screenshots of a UI and every changed region gets boxed and named: layout shift, spacing nudge, color restyle, text edit, visibility fade, added, or removed. Each box carries a confidence and the actual evidence (“the same content reappears 8px right, match 100%”). Pixel-diff tools already tell you where pixels differ; the useful part is naming the kind of change, so a reviewer can wave through nine color tweaks and stare hard at the one layout shift. Filter the overlay by type, then export an annotated PNG or a standalone HTML report for the PR.
The idea called for Claude Vision to do the classifying. It turned out to be classical image processing all the way down: a perceptual pixel diff with anti-aliasing detection finds the changes, clustering groups them into regions, and each region runs a gauntlet of checks (is one side just background, does the content reappear at an offset within 48px, is one side an alpha-blend of the other, do the edges correlate while the palette moved). Deterministic, offline, and screenshots never leave the browser. Heuristics have edges, though: a wide element nudged a few pixels produces two thin changed strips, and if they land far enough apart the tool reports an add plus a remove instead of one spacing change. 56 vitest tests, 96% coverage, no backend, no API keys.
See README for more details on the code or try it at https://ui-diff-lens.netlify.app
(Last update August 2026)
Shortcut Sprint
A flashcard app for your fingers. It shows a task (“Go to definition”, “delete the current line”), you press the shortcut, and SM-2 spaced repetition (the SuperMemo algorithm from 1990, same family as Anki) decides when you see that card again: instant recall pushes it out days, then weeks; fumbling brings it back tomorrow. VS Code, Chrome DevTools, Figma, and Vim come bundled, including multi-chord sequences like Ctrl+K Ctrl+S and Vim’s ciw, and instructors can upload a JSON set for whatever tool their course uses. A per-tool mastery radar and a daily streak provide just enough guilt to keep the habit going.
The idea called for Supabase to hold progress. I skipped it: localStorage does the job, so there are no accounts and nothing leaves the browser. The matching layer turned out to be the interesting part, because a keypress has two readings (the physical key and the character it produced) and Vim’s $ needs one while Ctrl+Shift+[ needs the other, so every press is matched both ways. The limitation is baked into the platform: a web page cannot capture browser-reserved shortcuts, Ctrl+W closes the tab no matter how much preventDefault you throw at it, so the one shortcut everyone knows is the one this app cannot teach. 78 vitest tests, 98% coverage, no backend, no API keys.
See README for more details on the code or try it at https://shortcut-sprint.netlify.app
(Last update August 2026)
Migration Diff Narrator
Paste two versions of a database schema, or two versions of your TypeScript interfaces, and get every change named and judged: safe, caution, or breaking, each with a one-sentence migration note (“Existing NULLs make this fail; backfill the column before adding the constraint”). Renames are inferred instead of being reported as a drop plus an add, so user_name becoming username shows up as a rename warning rather than a data-losing delete. One button copies the whole thing as a Markdown checklist ready for a PR description. Schema changes are the most dangerous part of a deploy, and the diff between two migration files is usually raw SQL a reviewer has to simulate in their head. This turns that into a ten-second scan.
The project idea called for the Claude API to classify each change. I skipped it. The taxonomy of schema changes is small and enumerable, so a fixed rule set does the job: same diff, same verdict, every time, offline, nothing you paste leaves the page. The honest limitation is parser coverage: it reads the common subset of SQL DDL (CREATE TABLE, ALTER, indexes, enums) and object-shape TypeScript interfaces, not every dialect corner, and it warns you when it skips something it does not understand. 74 vitest tests, no backend, no API key.
See README for more details on the code or try it at https://migration-diff-narrator.netlify.app
(Last update August 2026)
Schema Storyteller
Point it at a schema (SQL DDL, Prisma, or JSON Schema) and it tells you the story: which entities exist, how they relate, which tables are just join tables, and what usr_actv_flg was probably meant to say. A rule-based reviewer then lists what is likely missing: primary keys, indexes on foreign keys, uniqueness constraints. Export the whole narrative as Markdown and paste it into the design doc that should have existed in the first place.
No LLM here either, on purpose. A hand-written parser and a fixed rule set replace the suggested Claude API calls, so the narrative is free, offline, and reproducible, which also means it is templated: the prose will not win any literary awards, and the reviewer only knows the rules I taught it. 71 vitest tests, 88% coverage, fully client-side.
See README for more details on the code or try it at https://schema-storyteller.netlify.app
(Last update August 2026)
Pathfinding Playground
Draw a grid map, drop in walls and mud (mud costs 5 to enter), then watch A*, Dijkstra, BFS, DFS, and greedy best-first think one expansion at a time, each step narrated from the live algorithm state: which cell got picked, why, and what the g, h, and f values were. The project idea suggested Claude API calls for the narration. I skipped that. Every line is a template filled with the numbers the algorithm just computed, so narration is free, instant, works offline, and always matches what the search actually did, which an LLM cannot promise.
Each algorithm is a generator emitting a uniform event trace, and the canvas renderer and the narrator both consume the same events, so what you see and what you read cannot drift apart. Neighbor order and tie-breaking are fixed, so the same map always produces the identical trace; there is a test for that. The classroom payoff is compare mode plus share links: sprinkle mud, run BFS against Dijkstra side by side, watch BFS march straight through the expensive terrain while Dijkstra detours around it, then hand the exact puzzle to a class as a plain link, because the whole map, start, goal, and algorithm picks are run-length encoded into the URL hash. React + Vite, no backend, 20 vitest cases covering optimality, path validity, maze solvability, and trace determinism.
See README for more details on the code or play it at https://pathfinding-playground-pisanuw.netlify.app
(Last update August 2026)
Game Palette Inspector
Check a game palette against WCAG contrast and eight kinds of color vision deficiency, then get replacement colors that keep the art style. Roughly 8% of male players see color differently, and most game color tooling ignores them. Build a palette by hand, load a preset, or drop a screenshot and have its eight dominant colors pulled out; the vision lab then shows the palette, and the full frame, through protanopia, deuteranopia, tritanopia, their milder forms, and full and partial achromatopsia.
Two parts I like. The confusion report measures color pairs perceptually in OKLab and flags the ones that are clearly distinct with typical vision but collapse under a specific deficiency, naming the worst-case vision type. And the fix studio repairs a failing color by binary-searching OKLCH lightness with hue pinned, so the contrast target is met without repainting the art direction. It is also honest about impossibility: 7:1 against a mid-tone background is sometimes unreachable, and the tool says so instead of inventing a color. The simulation is Machado, Oliveira and Fernandes (2009), applied in linear sRGB and verified against the colour-science reference data. Everything runs in the browser: no accounts, no uploads, no tracking, no runtime dependencies beyond React, all the color math hand-rolled and unit tested. The interface follows its own advice, too: the accent is a protan/deutan-safe cyan and no pass/fail state is carried by color alone. Simulations are good approximations, not ground truth; individual perception varies.
See README for more details on the code or use it at https://game-palette-inspector.netlify.app
(Last update August 2026)
Google Flights MCP Server
An MCP server that hands Claude nine tools for searching flights and watching fares. Google has no official flights API, so this calls SerpAPI’s google_flights engine, which runs the search server-side and returns structured JSON. You can search one-way, round-trip, or multi-city, filter on cabin class, passengers, stops, and price, follow a departure_token to advance leg by leg through an itinerary, and follow a booking_token to see which providers sell the fare and at what price. “Nonstop round-trip SEA to NRT, Oct 3 back Oct 17, under $1200” is the whole interface.
The part I actually use is the price watching. Save a search as a watch with an optional target, and check_prices re-runs every watch, records what it finds, and reports drops and target hits, with a full price history per watch. A standalone checker script can be put on a schedule for hands-off alerts that fire a native macOS notification. Around 900 lines of Python, no build step, registers with Claude Code in a single command.
Honest limitation: this is scraping as a service, so it inherits SerpAPI’s rate limits (the free tier is 100 searches a month) and whatever Google changes about its result pages.
See README for setup
(Last update August 2026)
Teaching an LLM Tutor to Withhold the Answer
A paper on the engineering problem behind the Socratic guardrail: a capable model asked nicely, or asked angrily, will not reliably refuse to give an answer it could easily produce. A prompt is not enough. So the tutor enforces answer-withholding as a per-turn, machine-checkable contract instead. A non-LLM policy core reads only trusted learner state and sets a ceiling on an eight-rung help ladder, a deterministic detector strips solution code, and a separate LLM judge checks each risky reply against the contract before it is sent.
The tuning method may be the more portable contribution. Scripted student personas are driven through the live pipeline and re-scored by a stronger model, with every rejection’s stated reason recorded, so failures get fixed by cause rather than by vibes and no human subjects are needed to iterate. That surfaced an “over-help ladder” that I did not expect to be so orderly: fix the blatant solution leaks and the tutor starts naming the exact bug; fix that and it starts over-citing general facts. Each fix exposes the next rung. The measure, diagnose, and fix loop generalizes to any agent that has to refuse a capability it has.
Claude was a heavy collaborator on this one, which is a slightly recursive situation: a language model helping write up how to stop a language model from being too helpful.
The tutor it describes is YPDSA, deployed for CSS 342/343 and listed above.
Read it at https://arxiv.org/abs/2608.12292
(Last update August 2026)
Teaching Intro AI When the Tools Can Do the Homework
An experience report on redesigning CSS 382, our intro AI course, once it became clear that a language model could complete most of the assignments in it. The response was not to freeze the curriculum. The classical core stayed (search, adversarial search, MDPs, reinforcement learning) and a strand was added in which students build a language model from scratch, so that a tool they are required to use is also one they are required to understand. Assessment moved to work that resists unattributed automation: in-class exercises, reflective writing, and a defended team project. Examinations were removed entirely. The AI policy went from unmentioned in 2023 to required in 2026.
The center of the paper is the part I did not plan for. The cohort deliberated on and ratified a Student Bill of AI Rights governing my use of AI, including a provision that I personally complete any AI-generated assignment before issuing it. The prompt that scaffolded their deliberation was itself AI-generated, and the paper treats that provenance as part of the story rather than a footnote. It also reports the tensions, including students objecting to AI-generated course materials, and is explicit about how little a single-cohort design narrative can actually claim.
Read it at https://arxiv.org/abs/2608.05175
(Last update August 2026)
How of Happiness Quizzes
Ten chapter quizzes, twenty questions each, with Google sign-in and a leaderboard. React and Vite on the front, Supabase (Postgres, Auth, row-level security) on the back, Netlify for hosting.
Two design decisions carry the whole thing. First, your chapter score is the average of every attempt, not your best, which makes retakes free to take but not free to fail: open with 15 out of 20 and you are capped at 18.3 after three attempts and 19.5 after ten, approaching a perfect score without ever arriving. One bad attempt is permanent. Only a perfect first attempt scores perfectly. The global board sums your chapter averages, so breadth pays and every chapter you play can only add to your total. Second, players show as initials with no photo by default, because signing in with Google should not publish your full name and face to a public page. That default is enforced in the database rather than the interface: the leaderboard view returns a photo only when the player opted in, the profiles table is readable only by its owner, and players hold update rights on exactly two columns, so nobody can point their avatar at an arbitrary image.
Questions live in JSON, one file per chapter, with a seed script that validates before it writes.
See README for more details on the code or play it at https://how-of-happiness.netlify.app/
(Last update August 2026)
Building with AI: a mini workshop
A one-hour hands-on session that takes complete beginners from “AI is a mystery” to a working web app they built by describing it. Not a lecture about AI, and no code shown: the participants each end the hour with something on a screen that they made.
The repo is the whole kit rather than just the slides. A self-contained HTML deck, a participant handout with the prompt recipe, an easy/medium/hard menu of app ideas and ready-to-paste starter prompts, a list of a hundred small app ideas, and a facilitator guide with minute-by-minute timing, the live-demo script, rescue prompts for when someone is stuck, and troubleshooting. The handout is also translated into Turkish. Anyone can pick this up and run the session; that was the point of writing the facilitator guide rather than keeping the timing in my head.
Run it yourself from https://mini-ai-app-workshop.netlify.app
(Last update July 2026)
LTMS: A Truth Maintenance System in Python
A logic-based Truth Maintenance System and pattern-directed reasoning engine in pure Python, following Forbus and de Kleer’s Building Problem Solvers (MIT Press, 1993). The system maintains belief across a set of propositional clauses using Boolean Constraint Propagation, records well-founded support for every derived value, backtracks in a dependency-directed way when it hits a contradiction, and can explain why it believes anything it believes. Assert p or q, then assert not p, and it concludes q. Assume it is raining and it will conclude the ground is wet; retract the assumption and the wet ground goes back to unknown rather than lingering as an orphaned conclusion. That retraction behavior is the entire point of a TMS, and it is what separates it from a rule engine that only ever adds.
The reason to build it: JTMS and ATMS have a few toy ports, but the clausal-BCP LTMS with dependency-directed backtracking is close to unported outside Lisp and Racket. This is a clean, typed, tested version, roughly 6,300 lines across 16 modules, 143 tests, published on PyPI as ltms 0.1.0 under MIT. It goes past the base implementation into indirect proof, closed-world assumptions, dependency-directed search, prime implicates for logical completeness, and a SAT-style two-watched-literals BCP engine, with differential tests against PySAT. World models can live in .kb data files instead of Python, and those files carry expect lines that self-check when the file runs.
Alongside the code is a documentation site with a 17-chapter study companion walking through the concepts, the runnable examples, and worked solutions to the book’s exercises. The book itself comes from the Qualitative Reasoning Group at Northwestern, which is where my own qualitative-reasoning PhD came from, so this is partly a thirty-year round trip: the algorithms I learned in Lisp, rebuilt in Python with a language model as the pair programmer.
Still missing: a deployed site where you can type in rules and facts and watch belief change, retract an assumption, and ask the system why. That is the next piece.
See README for more details on the code, or read the study companion
(Last update August 2026)
Ladder Games
Two puzzle games in one app, sharing a structure and a stubborn design principle. Word Ladder asks you to change one letter at a time to climb from a start word to a target, scored on how close your path came to the shortest one. Number Ladder gives you a set of numbers and the operators + − × ÷ ^ and asks you to hit a target, Countdown-style, using each number at most once. Both offer Free Play and a Climb Mode that keeps raising the difficulty until you decide to stop.
The dictionaries are the interesting part of the word game: 26,419 entries across English words, Turkish words, English names, and Turkish names, at lengths 3 through 7. Each file holds only the largest connected component of the one-letter-change graph for that category and length, so puzzle generation is structurally incapable of handing out an unsolvable pair. The math side has the mirror-image guarantee: a memoized solver searches every reachable set of remaining numbers, both to verify a generated puzzle has a good solution and to show you the best possible result next to your own afterwards. The operator rules are deliberately unambiguous (subtraction is always |a − b|, division is always larger over smaller and only when it divides evenly) so a correct idea never fails on move order. UI is bilingual, English and Turkish.
The Hint button is the Socratic guardrail in miniature. Word Ladder tells you which letter position to change, never what to change it to. Number Ladder highlights which two numbers to combine, never the operator or the result. Each hint costs a flat 10 points off the round, so hint use limits itself without needing a cap. React + Vite, no backend required to play; an optional Supabase table adds a public leaderboard that records hints used alongside the score.
See README for more details on the code or play it at https://word-path.netlify.app/
(Last update August 2026)
Computing Power and Political Power
A UW faculty-led study abroad program in development for Winter 2028, co-directed with Asli Cansunar (Political Science, UW Seattle). Around twenty students spend the quarter in Istanbul taking three co-taught courses that pair a technical skill with a political-economic question: The Technology of Resistance (censorship, throttling, VPNs and Tor, measured with OONI, against the political science of networked collective action), Computational Political Economy of Turkey (OCR, geocoding, and choropleth mapping applied to Ottoman and Republican-era registers to see where public goods actually went), and a Fieldwork Practicum where mixed teams carry one applied project from question to public presentation. Computing students get the political economy, political science students get the command line, and every student does both. The site is four static pages on GitHub Pages with the three draft syllabi, the ten-week arcs, and a shared excursion week in Izmir and Ephesus.
The program is not an AI project, but it is an AI artifact: the site and the draft syllabi were written with Claude. The program still has to be reviewed by both instructors and confirmed by UW Study Abroad, so everything on the site is labeled a working draft and nothing is open for applications yet.
Read the drafts at https://pisanuw.github.io/turkey-study-abroad/
(Last update August 2026)
Emoji Lingua
Translate English into emoji, and emoji back into English. I love pizza on a rainy night becomes 👤 ❤️ 🍕 🔛 🌧️ 🌙; unknown words stay in place rather than disappearing, and a hint under the result tells you how many there were. The dictionary has 18,474 word-to-emoji entries (1,535 of them multi-word phrases) and 4,218 emoji-to-word glosses, generated from the Unicode CLDR annotations and then layered with hand-authored vocabulary: composed entries for abstractions (democracy → 🗳️, justice → ⚖️), function words that get dropped so the output reads cleanly, and curated overrides so the obvious choice wins (cat → 🐱, not 🐈). Phrases match greedily longest-first, so good morning → 🌅 rather than 👍 🌅, and a suffix-rule fallback catches forms the generator never materialized. The engine is hybrid: the dictionary is deterministic and always available, while Claude (when an API key is configured) handles context and can read an emoji sequence as a sentence instead of a word list. The AI path always degrades to the dictionary on any error, and the UI labels which engine produced each result, so a model outage cannot break the app. Express server wrapped as a Netlify serverless function, static page on the CDN, 44 tests at ~94% statement coverage. Honest limitation: emoji to English is a gloss, not grammar, so 🐱🍕 gives you cat pizza, not the cat ate pizza.
See README for more details on the code or use it at https://emoji-lingua-pisan.netlify.app
(Last update August 2026)
Canvas Group Evaluate
A command-line tool that generates a ready-to-deploy Google Form for peer and self evaluations from a Canvas group set. Two Python scripts do all the work: fetch_roster.py pulls groups and students from the Canvas API (or reads a CSV/TXT file you supply) and writes roster.json; generate_apps_script.py turns that into a self-contained Apps Script file you paste into script.google.com and run once. The resulting form uses section navigation to route each student to their own group’s rubric, with a configurable Likert block (Technical Contribution, Reliability, Communication, Problem Solving) plus two open-text questions per member. The form is locked to your Google Workspace domain (e.g. @uw.edu) so only signed-in students can respond. No web deployment — runs locally whenever the instructor needs to generate a new evaluation form.
See README for more details on the code.
(Last update May 2026)
Accessibility Lens
Paste any public URL and see it through four sets of eyes: low vision, color blindness, keyboard-only navigation, and screen reader. The app fetches the raw HTML, runs 11 WCAG 2.1 rules (missing alt text, broken heading structure, low color contrast, zoom disabled, duplicate IDs, unlabelled form controls, and more), and gives a concrete fix for each failure. The signature feature is a screen-reader view that strips all layout and shows the page as the linear stream of roles and text a blind user actually hears — making failures visceral rather than abstract. With an Anthropic API key, Claude rewrites each offending snippet into a corrected one; without it, rule-based fixes still cover every issue. Built as a TypeScript monorepo (React + Vite client, Express server), deployed on Render. 102 tests passing, ~96% server coverage. Built from the daily-project-ideas prompt for 2026-05-12.
See README for more details on the code or use it at https://accessibility-lens.onrender.com/
(Last update May 2026)
HTML Editor
A lightweight native macOS HTML editor built with SwiftUI, AppKit, and WebKit. Features a split-pane layout with a syntax-highlighted editor on the left and a live WKWebView preview on the right that updates automatically as you type. Includes a toolbar with one-click insertion of common tags (headings, bold, italic, links, lists), full file operations (New, Open, Save, Save As), and a status bar showing file path and cursor position. No web deployment — build and run from Xcode.
See README for more details on the code.
(Last update May 2026)
Digital Yusuf
A web app that lets you chat with a digital version of Yusuf Pisan. Built on a knowledge base of 10 structured persona files (~155K characters) covering bio, teaching philosophy, research, communication style, opinions, quirks, and full Substack articles — assembled from 30+ public sources including faculty pages, teaching evaluations, 63+ GitHub repos, Google Scholar, and promotion documents. The entire persona is loaded into the system prompt (no RAG needed at this scale). The digital Yusuf speaks in first person, matching his voice, humor, and opinions. Features streaming responses via SSE, optional Google Sign-In, BYOK (users can supply their own Anthropic API key), an admin dashboard with usage analytics and cost controls, and auto-email of finished conversations. Frontend is React + Vite on Netlify; backend is FastAPI on Render (Docker).
Chat with it at https://chatwithdigitalme.netlify.app/
(Last update May 2026)
Daily Project Ideas
Every morning at 5am, a Claude Code Remote Trigger wakes up, reads five subreddits (r/SideProject, r/sideprojects, r/ProductHunters, r/coolgithubprojects, r/AI_Agents), surfs the web, and commits three new project ideas to a public GitHub repo — prioritizing side projects, teaching tools, classroom assignments, and anything useful for students. By the time I sit down with coffee there is a fresh YYYY-MM-DD.md waiting. The repo also stores the full briefing, AI log, and change history so the prompt is version-controlled alongside the output. Read the Substack article My Cron Job Reads Reddit. The Prompts Are in Git. for the thinking behind it.
See README for more details on the code or browse the ideas at https://github.com/pisanuw/daily-project-ideas
(Last update May 2026)
Playful Interactions with AI
An opening keynote on how faculty can keep themselves up to date with AI tools and support their students. Covers hands-on examples of using LLMs for teaching, grading, and course content — with a warm, collegial tone aimed at fellow academics. Available in English and Turkish.
(May 2026)
Letter Game
A casual browser word game where you and the computer alternate naming things in a chosen category, one letter at a time from A to Z. Choose from 20 categories (Animals, World Cities, Movies, Mythological Deities, Dinosaurs, and more) or let the game pick at random. You type a word, it’s validated against the built-in word list, and the computer picks its own entry (with a Wikipedia image). You get 3 skips per game for tough letters. Built as a pure static site — plain HTML, CSS, and JavaScript with no framework or backend — using the Wikipedia REST API for images and Netlify Forms for word suggestions.
See README for more details on the code or use it at https://letter-category.netlify.app/
(Last update May 2026)
PasteMD
A Markdown publishing tool with passwordless authentication. Users sign in via magic link (email), write or paste Markdown, and publish it as a rendered page with a short shareable URL. Posts are stored in Netlify Blobs (no external database required). Authors can manage and delete their own posts from a personal dashboard; admins get a separate panel to view and delete any post across all users. Admins are notified by email whenever a new post is published. Built as a pure static frontend with Netlify serverless functions handling auth (JWT-based), post CRUD, and email delivery.
See README for more details on the code or use it at https://paste-md.netlify.app/
(Last update April 2026)
Grade Histogram Plotter
A lightweight Flask web app for visualizing grade distributions. Paste scores one per line (or upload a file), optionally customize the grade-bucket cutoffs, and get an instant histogram with summary statistics — mean, median, standard deviation, min, max, and per-bucket counts and percentages. Non-numeric entries are tallied in a separate NaN bucket. Includes CSRF protection and rate limiting. Deployed on Render.
See README for more details on the code or use it at https://grade-histogram-plotter.onrender.com/
(Last update April 2026)
UpvoteMe
A private, unlisted comment-and-voting app. Anyone can create a topic and instantly get two short URLs — one to share with participants, one private admin URL for moderation. Participants post comments (with optional file attachments, up to 5 files × 2 MB each) and upvote or downvote others. No public topic listing exists by design; topics are only reachable via shared URLs. Admins can lock or delete topics. Voting can be restricted to authenticated users (Google OAuth or magic link), in which case the admin can see who voted on each comment. Topics auto-delete after 30 days of inactivity.
See README for more details on the code or use it at https://upvoteme.netlify.app/
(Last update April 2026)
RankMe
A head-to-head voting app that ranks anything using the ELO rating system. Anyone can create a topic (movies, foods, photos, etc.), add items with optional images, and vote by repeatedly choosing between two randomly matched items. Rankings update in real time after each vote using a variable K-factor ELO algorithm (K=40 for new items, tapering to K=10 for established ones). Sessions are tracked via cookies to prevent repeat matchups in the same browser session. An admin panel at a secret URL allows deletion of topics and items.
See README for more details on the code or use it at https://rankme-1ttb.onrender.com/
(Last update April 2026)
The AI Grading Paradox
A classroom exercise for CSS 382 Introduction to Artificial Intelligence on how professors should grade homework in an era where AI can complete almost any assignment with near-perfect results. Groups of 5–6 students stress-tested four grading models (VIVA oral defense, GitHub audit, AI-hybrid rubric, and mastery/pass-fail), then designed their own “Ideal Grading Policy.” The assignment and a synthesis report of student submissions (created with Gemini) can be found here.
(Completed in April 2026)
The AI Learner’s Dilemma
A classroom exercise for CSS 382 Introduction to Artificial Intelligence on the ethical dilemma students face when using AI for homework: prioritize learning and risk a lower grade, or prioritize grades and miss the learning opportunity. Groups of 5–6 students analyzed three personas (The Accelerator, The Proxy, The Traditionalist), assessed long-term consequences through the lenses of technical interviews, employer expectations, and academic integrity, then drafted a 5-rule “Personal AI Ethics Protocol.” Near-unanimous consensus: The Accelerator is most sustainable, The Proxy is unethical, and a CS degree’s value shifts from code-writing to engineering judgment in a world of capable AI. The assignment and a synthesis report of student submissions (created with Gemini) can be found here.
(Completed in April 2026)
Ranked Voting
A full-stack ranked-choice voting web app. Admins create contests with multiple candidates, set a number of winners, control voter access via allowed email lists, and optionally randomize option order per voter. Voters submit drag-and-drop ballots. Results are computed using step-by-step Instant Runoff Voting (IRV), showing each elimination round until winner(s) are determined. Built with React + Vite, Tailwind CSS, Supabase (PostgreSQL + Auth), and Netlify serverless functions.
See README for more details on the code or use it at https://ranked-voting.netlify.app/login
(Last update April 2026)
40 Greatest Innovations – Ordering Game
A web-based card-ordering game where players arrange the 40 greatest innovations of all time in chronological order. Players drag or click cards across a deck, board, and “later” holding area, then press Check Answers to receive a final score (no hints during play). Supports drag-and-drop and touch.
See README for more details on the code or use it at https://order-cards.netlify.app/
(Last update April 2026)
ClaudeBot
Allows users to ask Claude questions, get Claude to review code, suggest patches via Discord. Users need to use their own Anthropic API key (so I do not have to pay for my students’ explorations!)
See README for more details. You can install it on your own Discord server using this link
(Last update April 2026)
MeetMe — Joint Meeting Finder
Allows gathering information from multiple people on their availability, so you can find a common meeting time. You can also set it up to let users “book” you based on your availability. Combination of https://www.when2meet.com/ https://calendly.com/ and https://doodle.com/
See README for more details on the code or use it at https://meetme.pisan.me/
(Last update April 2026)
Canvas Accessibility Fixer
Improves the accessibility score for Canvas pages automatically. Needs a canvas token to access Canvas pages.
See README for more details on the code or use it at https://canvas-accessibility.onrender.com/
(Last update April 2026)
TPS - Thermodynamics Problem Solver
Based on my PhD thesis, starting to recreate the project from scratch.
Use it at https://tps-thesis-recreation.onrender.com/
To be completed at a later date!
(Last update April 2026)
Grade Statistics
An internal project to generate graphs based on grades for each professor, for each course as well as historical data.
See README for more details on the code. Only used locally, so no web deployment.
(Last update April 2026)
Bullet Impact Simulator
A toy project on whether sensors could be used to determine where a bullet has hit on the target.
See README for more details on the code or use it at https://targetbullet.netlify.app/
(Last update April 2026)
Cloud Games Portal
Started with a simple version of Tic-Tac-Toe and ended up adding backgammon and chess as well. Nothing fancy, but was good for learning.
See README for more details on the code or use it at https://tic-tac-toe-app-857412880660.us-west1.run.app/
(Last update April 2026)
Choose Your Own Adventure
Starter code for student projects. Students extended this project to create Choose Your Own Adventure authoring tools as well as Choose Your Own Adventure reading platforms.
See README for more details on the code.
(Last update April 2026)
StockReptile - A Chess Playing Web App and Learning Tool
StockFish is the best chess engine out there. StockReptile is my inferior version but much more suited as a learning tool
See README for more details on the code or use it at https://stockreptile.netlify.app/
(Last update April 2026)
Teaching Evaluation Graph Generator
This tool automates the extraction of teaching evaluation metrics from historical course evaluation PDFs and visualizes them over time using line graphs.
See README for more details on the code or see my teaching evaluation dashboard at https://pisanorg.github.io/yusuf/graphs/
(Last update April 2026)
UW Course & Professor Finder
A web app from UW time schedule data so users can quickly answer:
- Which professor taught a specific course?
- Which courses did a specific professor teach?
- When (quarter and year) was the course offered?
The app supports all three UW campuses — Bothell, Tacoma, and Seattle — with a drill-down UI: select campus → select department → filter by course, professor, quarter, year.
See README for more details on the code or use it at https://uwcourses.netlify.app/
(Last update April 2026)
The Algorithmic Professor? AI Ethics in the Classroom
A classroom exercise on what students think of the professors’ use of AI in their teaching from generating slides to grading assignments with AI.
The assignment description and the summary report can be found here.
(Completed in April 2026)
CS Prof vs Gemini Prompts
Previously I had used Gemini to create slides. This time it claimed it cannot create slides. This is a summary of my “discussion” with Gemini trying to understand why it is refusing and trying to figure out a better prompt to get it to do what it had done before. You can read my “discussion” here
(Completed in March 2026)
History of AI slides by Gemini Pro
I asked (guided?) Gemini Pro to create a set of slides on the History of Artificial Intelligence that I plan to use in my CSS 382 Introduction to Artificial Intelligence course in Spring at University of Washington Bothell.
I provided History of Artificial Intelligence page from Wikipedia as a source. The slides produced by Gemini are here. I think it did a good job.
(Completed in March 2026)
| Yusuf Pisan | Computing & Software Systems (CSS) | University of Washington Bothell |