Casting for radio
Seven field runs testing whether a person will discover live radio by physically throwing a phone. The mechanism worked. The experiment that was supposed to prove people would enjoy it failed on its own terms, twice — and the tester asked to play again both times.
What was being tested
CatchFish turns radio discovery into fishing across Earth. You aim at a globe, load the rod by drawing the phone toward you, throw it away from you to cast, and reel a station in through sound and vibration. The catalogue is 31,557 real stations at their real coordinates.
Two questions had to be answered in that order. Can a phone browser hold a conveyor of live streams without the audio going wrong? Live radio has no seek, no buffer you control, and browsers block autoplay until a gesture. Then: does the waiting feel like suspense, or does it feel like buffering?
Thresholds for both were written down and frozen before any run, so a result could not be graded generously after the fact.
Phase A — the mechanism
Four runs on a Pixel 10 over cellular, 20 cast-and-reel cycles each, scored against the pre-registered bands.
| Run | Cohort | Caught | p50 | p95 | Verdict |
|---|---|---|---|---|---|
| 01 | Fresh origin | 17/20 | 2.29 s | 4.31 s | pass |
| 02 | Learned state | 19/20 | 2.00 s | 3.99 s | pass |
| 03 | Fresh, fixed rota | 17/20 | 1.83 s | 4.06 s | conditional |
| 04 | Learned, fixed rota | 18/20 | 1.67 s | 4.55 s | pass |
The primary question came back clean. Across 40 cycles and 79 candidate loads: zero stale or overlapping audio, zero post-unlock autoplay failures, zero stalls. All three audio elements unlocked from a single gesture and none was ever backgrounded. The conveyor works.
Two things the design could not measure
Runs 01 and 02 were built as a paired cohort: the same device and origin, once without learned state and once with. The second run was meant to show that remembering dead stations improves reliability. It could not have.
Run 02 drew none of run 01's 15 dead stations, and had almost no chance to. Fifteen dead entries in a catalogue of 31,557, over 40 random-bearing draws, gives 0.019 expected hits.
The second gap was the approach window — how late the station can be revealed before the wait stops working.
Run 05 — discarded
Thirty encounters were collected and then thrown out. Moving between screens in the test harness contaminated the thing being measured, so the run is marked discarded · harness-navigation-confound and scored nowhere. Its numbers are kept as diagnostics and excluded from every verdict.
Phase B — the feeling
Two complete, uninterrupted sessions. Twenty encounters each, four manual casts, every threshold as pre-registered.
| Pre-registered metric | Run 06 | Run 07 |
|---|---|---|
| Encounters rated | 40% fail | 100% pass |
| Waits felt as tension | 0% fail | 0% fail |
| Waits felt as buffering | 62.5% fail | 15% pass |
| Broken promises per 20 | 5 cond. | 7 fail |
| Reel rating, 1–5 | 4 pass | 5 pass |
| Finished without stopping | 20/20 pass | 20/20 pass |
| Would cast again | yes pass | yes pass |
| p50 time-to-first-sound | 1.75 s pass | 2.50 s pass |
| p95 time-to-first-sound | 4.82 s pass | 5.91 s pass |
| Cellular data per 20 | est. only no data | 6 MB pass |
| Stale or overlapping audio | 0 pass | 0 pass |
| Strike vibration | unfelt fail | 9/9 pass |
Run 07 is the cleaner failure, because almost everything else passed. Latency, measured Android data, audio isolation, static health and physical vibration all met their frozen bands. What failed was the registered mechanism itself: across 20 rated waits, not one was described as tension. Seven promised stations never arrived.
The loop works, by a different mechanism
The formal score does not describe what the tester did. They completed all four casts, rated the reel 5 out of 5, said they would cast again, and afterwards explained that what kept them winding was wanting to hear a different station — not suspense about whether one would arrive.
"Wind to find better station kept me going."
Run 06 supports the same reading from the other direction. Rating coverage was only 40%, but that was a design flaw in the controls, not indifference: the two questions asked about different moments, and all twelve unrated encounters were successful catches marked felt good. Thirteen of fifteen successful snaps felt good. Median listening time before rejecting a caught station was 9.72 seconds.
So the narrower supported conclusion, which is the one worth keeping:
A successful snap usually feels like a win. A miss can feel like buffering, and some successful stations still had a wait that felt like buffering.
That is a different product than the one that was registered. The registered bet was that the wait would carry the experience. The evidence says the rejection carries it — that people cast again because the last station was not the one they wanted, which makes an unbounded catalogue the asset and the suspense mechanic largely beside the point.
Status
Phase B is closed by product-owner decision. The Phase C build exists — a phone-first instrument using live location, compass bearing, and gravity-referenced raise-to-load and throw-to-cast detection — and its twenty-throw validation has not been run. The game is deferred; this page is the evidence, not the product.
Every figure here comes from run reports and raw event logs recorded at the time, in the CatchFish repository under tests/. Runs are numbered in collection order, and the discarded one keeps its number.