← Work

Casting for radio

Seven field runs testing whether a person will discover live radio by physically throwing a phone. The mechanism worked. The experiment that was supposed to prove people would enjoy it failed on its own terms, twice — and the tester asked to play again both times.

CatchFish Radio · Phases A and B Aug 2026 Pixel 10 / Android 17, cellular 150 encounters, 7 runs, 1 discarded

What was being tested

CatchFish turns radio discovery into fishing across Earth. You aim at a globe, load the rod by drawing the phone toward you, throw it away from you to cast, and reel a station in through sound and vibration. The catalogue is 31,557 real stations at their real coordinates.

The geometry of one cast A cone opens from the caster's position along a compass bearing of 57 degrees. Only stations inside the cone are eligible; the rest of the catalogue is out of reach for this cast. One station is locked as the target. you locked station cone, ±12° of 57°
One cast. The rod's compass bearing opens a cone, and only stations inside it are eligible — the rest of the 31,557 are out of reach until you turn and throw again. Bearings shown are the fixed rota the later runs fished: 57°, 78°, 144° and 201°, so both cohorts fished the same water.

Two questions had to be answered in that order. Can a phone browser hold a conveyor of live streams without the audio going wrong? Live radio has no seek, no buffer you control, and browsers block autoplay until a gesture. Then: does the waiting feel like suspense, or does it feel like buffering?

Thresholds for both were written down and frozen before any run, so a result could not be graded generously after the fact.

Phase A — the mechanism

Four runs on a Pixel 10 over cellular, 20 cast-and-reel cycles each, scored against the pre-registered bands.

RunCohortCaughtp50p95Verdict
01Fresh origin17/202.29 s4.31 spass
02Learned state19/202.00 s3.99 spass
03Fresh, fixed rota17/201.83 s4.06 sconditional
04Learned, fixed rota18/201.67 s4.55 spass

The primary question came back clean. Across 40 cycles and 79 candidate loads: zero stale or overlapping audio, zero post-unlock autoplay failures, zero stalls. All three audio elements unlocked from a single gesture and none was ever backgrounded. The conveyor works.

Two things the design could not measure

Runs 01 and 02 were built as a paired cohort: the same device and origin, once without learned state and once with. The second run was meant to show that remembering dead stations improves reliability. It could not have.

Run 02 drew none of run 01's 15 dead stations, and had almost no chance to. Fifteen dead entries in a catalogue of 31,557, over 40 random-bearing draws, gives 0.019 expected hits.

The 3 → 1 improvement in empty reels is not significant (Fisher exact, p = 0.605), and candidate-level failure was unchanged: 38.5% versus 37.5%, p = 1.000. The pair is valid evidence for timing and mechanism. It is not evidence about learned state.

The second gap was the approach window — how late the station can be revealed before the wait stops working.

Primary success by approach lead, with 95% intervals Four lead times were tested ten times each. Success is non-monotonic and every 95% interval overlaps every other, so ten trials per rung cannot determine where the window should open. 0% 25% 50% 75% 100% 6 s 8/10 4.5 s 8/10 3 s 5/10 2 s 7/10 primary success, with 95% Wilson intervals
Success is non-monotonic, and every interval overlaps every other one. The largest contrast reaches only p = 0.350. Ten trials per rung cannot locate the window, so it stays undetermined and is recorded that way rather than read off the highest dot.

Run 05 — discarded

Thirty encounters were collected and then thrown out. Moving between screens in the test harness contaminated the thing being measured, so the run is marked discarded · harness-navigation-confound and scored nowhere. Its numbers are kept as diagnostics and excluded from every verdict.

Phase B — the feeling

Two complete, uninterrupted sessions. Twenty encounters each, four manual casts, every threshold as pre-registered.

Pre-registered metricRun 06Run 07
Encounters rated40% fail100% pass
Waits felt as tension0% fail0% fail
Waits felt as buffering62.5% fail15% pass
Broken promises per 205 cond.7 fail
Reel rating, 1–54 pass5 pass
Finished without stopping20/20 pass20/20 pass
Would cast againyes passyes pass
p50 time-to-first-sound1.75 s pass2.50 s pass
p95 time-to-first-sound4.82 s pass5.91 s pass
Cellular data per 20est. only no data6 MB pass
Stale or overlapping audio0 pass0 pass
Strike vibrationunfelt fail9/9 pass
Formal Phase B verdict: FAIL. Both runs. The verdicts are retained exactly as pre-registered.

Run 07 is the cleaner failure, because almost everything else passed. Latency, measured Android data, audio isolation, static health and physical vibration all met their frozen bands. What failed was the registered mechanism itself: across 20 rated waits, not one was described as tension. Seven promised stations never arrived.

The loop works, by a different mechanism

The formal score does not describe what the tester did. They completed all four casts, rated the reel 5 out of 5, said they would cast again, and afterwards explained that what kept them winding was wanting to hear a different station — not suspense about whether one would arrive.

"Wind to find better station kept me going."

Run 06 supports the same reading from the other direction. Rating coverage was only 40%, but that was a design flaw in the controls, not indifference: the two questions asked about different moments, and all twelve unrated encounters were successful catches marked felt good. Thirteen of fifteen successful snaps felt good. Median listening time before rejecting a caught station was 9.72 seconds.

So the narrower supported conclusion, which is the one worth keeping:

A successful snap usually feels like a win. A miss can feel like buffering, and some successful stations still had a wait that felt like buffering.

That is a different product than the one that was registered. The registered bet was that the wait would carry the experience. The evidence says the rejection carries it — that people cast again because the last station was not the one they wanted, which makes an unbounded catalogue the asset and the suspense mechanic largely beside the point.


Status

Phase B is closed by product-owner decision. The Phase C build exists — a phone-first instrument using live location, compass bearing, and gravity-referenced raise-to-load and throw-to-cast detection — and its twenty-throw validation has not been run. The game is deferred; this page is the evidence, not the product.

Every figure here comes from run reports and raw event logs recorded at the time, in the CatchFish repository under tests/. Runs are numbered in collection order, and the discarded one keeps its number.