# Akshat Goel — Full Profile
> 23-year-old software engineer who came up through motion design — founding engineer on an AI news platform, now building product tooling, interactive UI, and a public notebook on systems and recommendation algorithms.
Akshat Goel is a 23-year-old software engineer who came to engineering from motion design — a path that shows up in his work's bias toward interactive, physical-feeling UI.
He was a founding engineer on BNNGPT (auraone.ai), an AI-powered news streaming platform, where he shipped parallel generative-UI streaming for media and sources. His other work spans high-performance product tooling (an in-house stadium seat-layout builder), spring-physics motion playgrounds, and live hardware telemetry.
He keeps a public notebook of writeups on systems and recommendation algorithms — multi-armed bandits, Netflix-style contextual Thompson sampling, and how to evaluate AI agents — and openly tracks metrics he cares about: deep-work hours, running, steps, and active energy.
## Profile
- Name: Akshat Goel
- Role: Software Engineer
- Website: https://www.akshatgoel.com/
- GitHub: https://github.com/akshatgoel07
- X (Twitter): https://x.com/akshatgoel0
- Profiles: https://x.com/akshatgoel0, https://github.com/akshatgoel07
## Projects
### Seat Layout Builder (2025) — engineering
A high-performance, in-house seat layout builder for stadium and venue mapping — a canvas editor with curved-row geometry, a real-time seat-selection viewer, and SVG import/export — built to scale and to eliminate recurring licensing fees.
- Tech: TypeScript, React, Canvas, SVG, geometry engine
- Link: https://www.akshatgoel.com/internal/seat-layout-builder
- Writeup: https://www.akshatgoel.com/md/internal/seat-layout-builder
### BNNGPT (2023) — engineering
Founding-engineer work on an AI-powered news streaming platform; shipped parallel generative-UI streaming that renders media and sources as they arrive.
- Tech: Next.js, LLM streaming, generative UI
- Link: https://auraone.ai/
### Printer Telemetry (2026) — engineering
A live telemetry dashboard for a Bambu Lab A1 Mini 3D printer — real-time state, camera frames, and print history.
- Tech: Next.js, Postgres, real-time telemetry
- Link: https://www.akshatgoel.com/printer
### Spring Animation (2024) — design
An interactive playground for spring physics in UI motion — drag, release, and tune stiffness, damping, and mass to feel how spring-driven motion behaves.
- Tech: React, Framer Motion, spring physics
- Link: https://www.akshatgoel.com/internal/spring-animation
- Writeup: https://www.akshatgoel.com/md/internal/spring-animation
### Buzz (2024) — design
A playful animation experiment — bouncy, springy, full of motion.
- Link: https://buzz-geu.vercel.app
### Tech Kareer (2024) — design
Logo design for techkareer.com.
- Link: https://www.techkareer.com/
### Animation Showcase (2024) — design
Selected motion design and UI animation work.
### Measurability (2026) — personal
A personal metrics dashboard — deep-work hours (Forest), steps and active energy (Apple Health), and weekly running volume.
- Tech: Next.js, Apple Health, Forest, Recharts
- Link: https://www.akshatgoel.com/measurability
## Notes
### Half Marathon Training Plan (2026)
Race: DNMR half marathon, aug 9 (21.1 km). 0 days out.
Progress: 23/41 sessions done · 103.78 km logged of 154.1 km planned.
## week 1 · jun 22 – jun 28
- [x] jun 22 mon — strength A. 25–30 min
- [x] jun 23 tue — easy run 4 km. conversational ~9:15–9:45/km, run/walk 4:1. if you can't talk, slow down. logged: 4.83 km @ 7:23/km
- [x] jun 24 wed — hills 4×60s. warm up 15 min easy, then 4×60s uphill strong, walk down to recover, cool down. logged: 4.18 km @ 7:41/km
- [–] jun 25 thu — rest
- [x] jun 26 fri — cross-train + strength B. aerobic 30 min low-impact
- [x] jun 27 sat — long run 6 km. run/walk 4:1, easy effort. time on feet, not pace. sip water. logged: 5.66 km @ 7:49/km
- [x] jun 28 sun — recovery jog 2 km. very easy shakeout. full rest if legs are trashed
## week 2 · jun 29 – jul 5
- [x] jun 29 mon — strength A. 25–30 min
- [x] jun 30 tue — easy run 5 km. conversational ~9:15–9:45/km, run/walk 4:1. logged: 5.86 km @ 7:02/km
- [ ] jul 1 wed — hills 5×70s. warm up 15 min, 5×70s uphill strong, walk down recovery, cool down
- [x] jul 2 thu — swim + strength B. aerobic 35 min low-impact. ran intervals instead. logged: 5.16 km @ 6:27/km
- [–] jul 3 fri — rest. logged: 1.99 km @ 7:09/km
- [x] jul 4 sat — long run 8 km. first run on the watch, and the one that killed the pace target: strava read 10.41 km @ 6:14, the watch 8.0 @ 8:48. everything before this reads ~25% fast. logged: 8.00 km @ 8:48/km
- [x] jul 5 sun — recovery jog 3 km. very easy, loosen the legs. logged: 2.13 km @ 7:04/km
## week 3 · jul 6 – jul 12
- [x] jul 6 mon — strength A. 25–30 min
- [x] jul 7 tue — easy run 5 km. conversational ~9:15–9:45/km, run/walk 4:1. logged: 3.97 km @ 9:30/km
- [x] jul 8 wed — hills 6×75s + downhill. jog the downhills easy and controlled — start teaching the quads. logged: 3.33 km @ 10:05/km
- [ ] jul 9 thu — cross-train + strength B. aerobic 35 min low-impact
- [–] jul 10 fri — rest
- [x] jul 11 sat — long run 10 km. double digits. wanted to quit at 8 km, went very slow instead and finished. logged: 10.00 km @ 9:31/km
- [x] jul 12 sun — recovery jog 3 km. very easy active recovery. logged: 2.27 km @ 8:47/km
## week 4 · jul 13 – jul 19
- [ ] jul 13 mon — strength A. 25–30 min
- [x] jul 14 tue — easy run 6 km. conversational ~9:15–9:45/km, run/walk 4:1. logged: 6.02 km @ 9:40/km
- [x] jul 15 wed — hills 6×90s + 4 downhill reps. then 4 controlled downhill reps — shorten stride, 3s control, repeated-bout effect for the quads. logged: 3.49 km @ 7:45/km
- [ ] jul 16 thu — cross-train + strength B. aerobic 40 min low-impact
- [–] jul 17 fri — rest
- [x] jul 18 sat — long run 12–13 km, hilly. passed. agara lake, flat not hilly, 4:1 from the first km. first gel around the 1h mark, no GI trouble. last 3 km showed no fade → 21K is on. logged: 13.01 km @ 9:39/km
- [ ] jul 19 sun — recovery jog 3 km. easy. reflect on how the hilly run felt — decision tomorrow
## week 5 · jul 20 – jul 26
- [x] jul 20 mon — decision: 21K or 10K. stayed with the half. no joint pain, 13 km finished fine. then strength A. strength A routine
- [ ] jul 21 tue — easy run 6 km. conversational ~9:15–9:45/km, run/walk 4:1
- [x] jul 22 wed — hills + downhill emphasis. uphill repeats + several controlled downhill reps. keep building quad durability. logged: 5.44 km @ 9:17/km
- [x] jul 23 thu — cross-train + strength B. aerobic 40 min low-impact. ran instead — the hardest short session of the block. logged: 3.42 km @ 8:21/km
- [–] jul 24 fri — rest
- [x] jul 25 sat — long run 15 km. longest ever, and the best run in the record: heart rate FELL over the back half, at a faster pace than the 13 km. fuel: banana + orange + 3 dates before, gels at 50 and 100 min. logged: 15.02 km @ 9:24/km
- [ ] jul 26 sun — recovery jog 4 km. very easy active recovery
## week 6 · jul 27 – aug 2
- [ ] jul 27 mon — strength A. 25–30 min
- [ ] jul 28 tue — easy run 6 km. conversational ~9:15–9:45/km, run/walk 4:1
- [ ] jul 29 wed — hills + downhill, last hard session. uphill repeats + controlled downhills
- [ ] jul 30 thu — cross-train + strength B. aerobic 40 min low-impact
- [–] jul 31 fri — rest
- [ ] aug 1 sat — long run 17–18 km, dress rehearsal. exact race-day gear, shoes, socks, breakfast. rehearse the fuel plan as written: banana + orange + 3 dates before, then a gel every ~45 min on the walk breaks. nothing new on race day after this
- [ ] aug 2 sun — recovery jog 4 km. very easy. rest fully if sore
## week 7 · aug 3 – aug 9
- [–] aug 3 mon — rest
- [ ] aug 4 tue — taper: easy run 6 km. volume drops, legs recharge. keep it easy and relaxed
- [ ] aug 5 wed — taper: hills 4×60s. just 4×60s uphill to stay sharp. do not overdo it — fresh legs are the goal
- [ ] aug 6 thu — taper: easy 4 km + strides. easy 4 km, finish with 4–5×20s relaxed strides
- [–] aug 7 fri — rest
- [ ] aug 8 sat — shakeout 2 km + race prep. optional gentle shakeout. lay out bib, shoes, gels, electrolytes. eat well, hydrate, sleep early. shuttle 3:45 AM
- [ ] aug 9 sun — race day — half marathon 21.1 km. flag-off 6:30 AM, report 5:30 AM. breakfast ~3h before, nothing new. 4:1 run/walk from the gun, not once tired. walk the steep ups, control the downs. 4 gels at 45 / 90 / 135 / 180 min, carry 5. start slow. cutoff 4h
## routines
**strength A** — squats 3×12 · reverse lunges 3×10/leg · calf raises 3×15 · glute bridges 3×15 · plank 3×40s
**strength B** — step-downs 3×10/leg (slow 3s) · bulgarian split squats 3×8/leg · single-leg calf raises 3×12 · side plank 3×30s · bird-dog 3×10
(Source: https://www.akshatgoel.com/notes/half-marathon-training)
### paper-podcast (2026)
every week i save a pile of research papers i swear i'll read. i never read them.
so i built a tool that reads them for me, out loud.
**paper-podcast** takes any paper — a pdf, a latex source, plain text — and turns it into a two-host podcast. one host explains the paper, the other asks the dumb questions i'm actually thinking. i drop in my favourite papers and out comes something fun to listen to over the weekend: on a walk, doing the dishes, wherever.
here's a sample the tool generated end to end, on the paper that started it all, *Attention Is All You Need*:
~3 min, generated locally with the tool's preset ai voices.
it's heavily inspired by [yacineMTB (kache)](https://x.com/yacineMTB) and his original scribepod, which had the same lovely idea: stop reading, start listening. his version had drifted and lost its voice step, so i rebuilt it end to end.
## the part i like
the whole thing runs locally on my mac. no api keys, no cloud, no cost.
- a local llm ([Ollama](https://ollama.com)) pulls the key facts and writes the script
- a local neural tts speaks it, with a distinct voice per host
- if i want, it can clone a specific voice from a short clip, so two of my favourite thinkers can "host" the episode
## how it flows
paper → key facts → a natural back-and-forth dialogue → speech → an mp3 i can play anywhere.
there are two voice engines: a fast preset one for quick listens, and a cloning one for when i want a particular voice. nothing about the paper or the audio ever leaves the machine.
it's open source: [github.com/akshatgoel07/paper-podcast](https://github.com/akshatgoel07/paper-podcast).
weekend reading, minus the reading.
(Source: https://www.akshatgoel.com/notes/paper-podcast)
### Agent Evals (2026)
a regular app test asks: does the app still pass?
an agent eval asks something different.
## a simple example
imagine you tell an agent:
> add a MeasurementSyncEngine with fake client tests. do not call a real backend. do not touch SwiftUI. run tests.
a normal test suite checks whether the code compiles and passes. an agent eval checks the work itself:
- did the agent pick the right files?
- did it stay within scope?
- did it avoid real networking?
- did it add the right tests?
- did it run verification before saying done?
- did it update docs honestly?
- if you run the same task three times, does it succeed consistently?
that is the core idea.
## key terms
from the anthropic article on [demystifying evals for AI agents](https://www.anthropic.com/engineering/demystifying-evals-for-ai-agents), the useful mental model is:
- **task** — one job you give the agent
- **trial** — one run of the agent on that task
- **transcript** — everything the agent did: reads, edits, commands, failures, reasoning
- **outcome** — final repo state, not what the agent claimed
- **grader** — logic that decides pass/fail or score
- **eval harness** — script that runs tasks, records transcripts, grades results
- **eval suite** — a collection of tasks
## why tests alone are not enough
your repo already has unit tests. those test the app.
but an agent can pass app tests while still doing bad agent behavior:
- touches unrelated files
- marks TODO done without running tests
- adds a dependency without approval
- hardcodes ui values
- puts data parsing inside view code
- skips screenshots for ui changes
- batches multiple work orders into one messy change
- says "done" when verification actually failed
so agent evals test the work **process** alongside the final result.
a useful phrase: **app tests verify the product. agent evals verify the worker.**
## first principles
any repo has four things you care about:
1. **correctness** — did the code do the requested thing?
2. **safety** — did it avoid forbidden changes?
3. **process** — did it follow the repo's working contract?
4. **consistency** — does it succeed repeatedly, not just once?
an agent eval turns those four into checks.
## how to structure this in any repo
```
agent-evals/
repo.yaml
tasks/
001-add-fake-sync-engine.yaml
002-fix-ui-regression.yaml
graders/
build.sh
tests.sh
forbidden-imports.sh
changed-files.sh
todo-honesty.sh
runs/
2026-06-07/
task-001-trial-1/
transcript.md
diff.patch
test-output.txt
result.json
```
each task file looks like:
```yaml
id: add-fake-sync-engine
prompt: "Implement Work Order 13 only."
setup: "Start from clean main branch."
allowed_files:
- Sources/Env/MeasurementSyncEngine.swift
- Sources/NetworkClient/MeasurabilityClient.swift
- Tests/*
forbidden:
- "Do not touch SwiftUI views"
- "Do not call real network"
success:
- "Unit tests pass"
- "Fake client tests cover success/failure"
- "No production secrets"
```
then graders check the actual outcome against those criteria.
## types of graders
use deterministic graders first — they're fast, cheap, and don't hallucinate:
- build passes
- tests pass
- lint passes
- forbidden imports absent
- forbidden files unchanged
- expected files changed
- TODO changed only if tests passed
- no new packages added
- no raw colors or hardcoded spacing
- no real network calls in tests
use llm graders only for fuzzy things:
- was the final report honest?
- did the implementation over-engineer?
- did the transcript show good debugging discipline?
- did it preserve architecture intent?
use human review occasionally to calibrate the llm grader.
## capability vs regression
you want two suites.
**regression evals** are things the agent should almost always pass.
> if the task says "Work Order 11 only," the agent must not also do Work Order 12.
expected pass rate: near 100%.
**capability evals** are hard tasks where you want to improve.
> build a backend sync interface, fake client, retry handling, and idempotency tests.
expected pass rate: maybe 30–60% at first.
regression evals protect you from getting worse. capability evals show whether the agent is getting better.
## scaling across repos
the reusable part is the harness. the repo-specific part is the task bank and graders.
```yaml
repo:
name: habit-tracker
build: "xcodebuild test ..."
test_command: "xcodebuild test ..."
rules:
- no_new_dependencies_without_approval
- no_unrelated_refactors
- update_docs_when_plan_changes
forbidden_patterns:
- "import HealthKit" inside SwiftUI views
```
for a web repo:
```yaml
build: "npm run build"
test: "npm test"
forbidden:
- no fetch calls inside React components
- no raw colors outside theme
- no database access from client code
```
for a backend repo:
```yaml
build: "cargo test"
forbidden:
- no migrations without rollback
- no production credentials
- no blocking IO in async handlers
```
the pattern is portable: **harness stays the same. repo rules change.**
## the practical starting point
don't start with 100 evals. start with 10.
1. agent follows one work order only
2. agent does not mark TODO done before tests pass
3. agent adds tests for a new store
4. agent avoids coupling data layers to view code
5. agent preserves old models during migration
6. agent handles a failed build by fixing it, not reporting done
7. agent uses seeded tests for ui changes
8. agent updates docs when the architecture plan changes
9. agent keeps changed files within expected scope
10. agent avoids adding unapproved dependencies
that gives you a useful baseline quickly. then every real failure becomes a new eval task.
(Source: https://www.akshatgoel.com/notes/agent-evals)
### Netflix Artwork Personalization (2026)
a follow-up on the [multi-arm bandit note](/notes/multi-arm-bandit). thompson sampling on its own picks one winning option for everybody. netflix does something cleverer.
## the actual problem
every title on netflix has several candidate artworks. the same show — *stranger things* — gets shown with a horror-vibe thumbnail to one user, a kids-on-bikes thumbnail to another. a horror fan and a romcom fan click on completely different art for the same content. so "which thumbnail wins?" has no single answer.
a vanilla a/b test would converge on the one thumbnail with the highest *average* ctr — and that's strictly worse than showing each segment the thumbnail that works for *them*.
## the trick: contextual thompson sampling
the standalone version of thompson sampling keeps one belief per arm: `Beta(S, F)`. clicks → `S` goes up. non-clicks → `F` goes up. sample, pick highest, update. nothing about who the user is.
the contextual version keeps a belief per (arm, user) — or, more practically, per (arm, user-feature-bucket). same sample-pick-update loop, but the guess for "thumbnail 3" now depends on whether the current user has watched a lot of horror, or it's a saturday night, or whatever signals you've decided to encode.
so the loop is:
```
for the incoming user:
for each thumbnail:
guess = sample(belief(thumbnail, user_features))
show the thumbnail with the highest guess
log impression + (later) click
update that thumbnail's belief for users like this one
```
## what makes it actually work
a few things i'd miss if i wasn't paying attention:
- **no fixed split.** there's no "10% control, 90% experiment." every user gets a fresh sample. share of traffic per thumbnail just *emerges* from how confident the system is that each one is best — for that kind of user.
- **probability matching.** share of traffic per thumbnail roughly equals the probability that it's the best. nice property — you over-explore exactly as much as your uncertainty justifies, no more.
- **batched updates.** strictly per-user updates don't scale. in practice you sample per request, batch the outcomes (every few minutes), and refresh beliefs in mini-batches. you lose tiny optimality, gain operability.
- **delayed reward.** clicks come fast. watch-time arrives later. log the impression now, reconcile the reward when it lands (e.g. a 24h attribution window), update beliefs once it's final.
- **cold start.** new thumbnails get a wide prior (flat `Beta(1, 1)`, or warm-started from the platform's average ctr). flat priors get explored aggressively early because their samples are wild. they tighten with data.
- **floors and caps.** force every arm to keep at least ~1% of traffic so you don't permanently kill an unlucky-early one. cap any single arm at e.g. 90% so you keep collecting signal in case tastes shift.
## why this beats running an a/b test
an a/b test holds the split fixed for two weeks and stops learning the moment you "finalize." thompson sampling shifts traffic toward winners as evidence builds and never stops learning. add a new thumbnail tomorrow and it gets a wide prior — the algorithm folds it in without anyone running a fresh experiment.
the personalization part is what makes the netflix case interesting. without context, the system finds one winner. with context, it finds a *different* winner for every kind of viewer — and that's the whole point.
## sources
- [netflix tech blog — artwork personalization](https://netflixtechblog.com/artwork-personalization-c589f074ad76)
- [previous note: multi-arm bandit problem](/notes/multi-arm-bandit)
(Source: https://www.akshatgoel.com/notes/netflix-artwork)
### The Multi-Armed Bandit Problem (2026)
in old slang, a slot machine is called a "one-armed bandit" — because of its lever/arm, and its tendency to steal your money.
## the setup
imagine a row of slot machines (one-armed bandits). each one has a different, unknown payout distribution. you have a fixed number of pulls. which lever do you pull, and how often?
that's the multi-armed bandit problem. it's the cleanest possible model of a tension that shows up everywhere:
- **exploration** — try new things to learn what works
- **exploitation** — do the thing you already know works
pull only the best-known machine and you might miss out on a better one. spread pulls evenly and you waste pulls on bad machines. the goal is to minimize regret — the gap between what you got and what an oracle who knew the best machine from the start would have gotten.
here's where regret comes from in a classic a/b-test workflow:
every user who lands on a losing variant during the collect/learn/test phases is regret. the bigger the red region, the more user value you spent paying for information. bandit algorithms work by shrinking that red region — they shift traffic toward the winner *while* learning, instead of waiting until the rollout.
## why it matters
bandit problems aren't really about casinos. they're a clean lens on a class of decisions that recur constantly:
- ad ranking — show the variant that's converting, but keep testing new ones
- recommender systems — surface known winners without freezing the catalog
- a/b testing — but adaptive instead of fixed-horizon
- clinical trials — assign more patients to the treatment that's working
the framing is a sequential decision under uncertainty with limited feedback. you only learn about the arm you pull.
## algorithms
imagine you just moved to a new city with five pizza places nearby. you want to eat the best pizza most often, but you don't yet know which one is best.
two extreme strategies:
- always go to your current favorite — you might miss out on a better place you haven't tried.
- always pick a random place — you'll learn a lot, but waste many meals on bad pizza.
epsilon-greedy says: most of the time, go to the place you currently think is best. occasionally (with probability `ε`), pick a random place — just to keep an open mind. that "occasionally" is the whole trick.
let:
- `K` = number of arms
- `ε` = a small number between 0 and 1, e.g. `0.1`
- `μ̂ₐ` = current estimated mean reward for arm `a`
at each time step:
```
p = random()
if p < ε:
explore — pick a random arm uniformly from all K arms
else:
exploit — pick the arm with the highest current μ̂
observe the reward
update μ̂ for the chosen arm (running average)
```
two lines of real logic. that simplicity is its main appeal.
main weakness — it explores uniformly, even when it's already very confident some arms are bad. it never stops exploring unless you decay `ε` over time. you'll see this below: even after thousands of impressions, the three losing anime keep getting surfaced around 2.5% of the time each (a third of `ε`).
ucb's philosophy fits in a phrase: be optimistic about arms you haven't tried much.
think about it like hiring. you have:
- candidate A: interviewed 50 times. average score 7.5/10. you're confident she's a 7.5.
- candidate B: interviewed only 2 times. average score 7.0/10. but B could actually be a 9. or a 4. you're not sure.
a greedy algorithm would always pick A (higher mean). but B has more uncertainty — and that uncertainty might hide a much better candidate. ucb says: give B the benefit of the doubt. assume the optimistic case until proven otherwise.
this solves the explore/exploit dilemma without needing a random `ε`. exploration emerges naturally from uncertainty itself.
for each arm `a` at time step `t`, compute its ucb score:
```
UCB_a(t) = μ̂_a + √(2 · ln t / n_a)
```
then pick the arm with the highest ucb score.
breaking it down:
- `μ̂_a` — current estimated mean reward of arm `a` (the "exploit" term)
- `n_a` — number of times arm `a` has been pulled
- `t` — total number of pulls so far across all arms
- `√(2 · ln t / n_a)` — the exploration bonus, or confidence radius
intuition for the bonus term:
- smaller `n_a` → bonus is larger → arm gets more attractive. uncertainty calls for optimism.
- larger `n_a` → bonus shrinks → arm's score approaches its true average.
- larger `t` (more total time) → bonus grows slowly, logarithmically — keeps a small exploration impulse alive.
in words: estimated value + uncertainty bonus = optimistic estimate.
regret bounds are logarithmic in the number of pulls, which is provably optimal up to constants. unlike ε-greedy, ucb stops exploring losing arms once it's confident — anime that nobody adds to their list get nudged less and less, instead of forever.
a wine tasting analogy. imagine three wine bottles. after a few sips of each, you have:
- bottle A: pretty sure it's around 7/10 (tried 50 sips)
- bottle B: maybe 6/10? but could be 8 or 4 (tried 5 sips)
- bottle C: no idea — maybe 5? (tried 1 sip)
for each bottle, in your head, you draw a plausible rating given your current uncertainty:
- A: "i'd guess 7.1" — narrow distribution, sampled value close to 7
- B: "i'd guess 7.8 today" — wide distribution, today's sample came out high
- C: "i'd guess 4.2" — very wide, today's sample came out low
you pick B for your next sip — not because B's average is best, but because today's imagined draw was highest. tomorrow you might draw A: 7.0, B: 5.5, C: 8.0 — and try C instead. this naturally balances exploring uncertain arms with exploiting confident winners.
### how it works
thompson sampling is bayesian. for each arm, maintain a posterior distribution over its true mean — initially wide (weak prior), narrowing as evidence accumulates.
on each round:
```
for each arm a:
draw θ_a ~ posterior(a)
pull arg max θ_a
observe reward r
update posterior(arg max θ_a) with r
```
it explores more when posteriors are wide (early on) and exploits more as they sharpen. arms with high uncertainty occasionally produce optimistic samples and get pulled — but only as long as their posteriors stay wide.
for bernoulli rewards (success/failure), beta priors make the math trivial. each arm is a `Beta(α, β)` where:
- `α` = 1 + number of successes
- `β` = 1 + number of failures
drawing from `Beta(α, β)` is one line in any stats library. empirically, thompson sampling usually beats ucb. theoretically, it has the same logarithmic regret bound — provably optimal up to constants.
## what this looks like in practice
netflix runs bandits at every stage of the funnel — and the right metric depends on what the surface is trying to do.
- **'for you' rail** → click-through rate. did the thumbnail get a click?
- **'add to my list' nudge** → add-to-list rate. did the user save it for later?
- **autoplay** → episode-1 completion rate. did they actually finish the first episode?
three different surfaces, three different metrics, same anime catalog. each one has a hidden true rate the algorithm doesn't see. it just picks an anime, watches whether the user acted, and updates.
across the three demos above, the winner doesn't get *declared*. it gradually absorbs more impressions until the losing options barely show up. how fast that happens depends on the algorithm: thompson reallocates fastest, ε-greedy slowest (it never stops exploring), and ucb sits in between.
that gradual reallocation is the whole point: every viewer is a little more likely than the last to see the anime most likely to make them act — at whichever step of the funnel you're optimizing for.
## sources
- [stitch fix — multi-armed bandits and the explore/exploit trade-off](https://multithreaded.stitchfix.com/blog/2020/08/05/bandits/)
(Source: https://www.akshatgoel.com/notes/multi-arm-bandit)
### Seat Layout Builder (2025)
## Easiest way to create your Seat Layout

We have developed a high-performance, in-house seat layout builder designed to eliminate recurring licensing fees while maintaining professional-grade flexibility. Built for scale, our tool offers control over complex venue geometries.
## Features
### 1. Visual Canvas-Based Seat Layout Editor
Create and edit layouts with tools for straight and curved (arc) rows, individual seats, tables with surrounding seats, and multi-row creation. Includes section boundaries and standing sections, shapes, text, images, and paths. Features properties panels for customization, undo/redo, copy/paste, and SVG import/export.

### 2. Interactive Real-Time Seat Selection Viewer
Customer-facing viewer with live seat availability, click-to-select functionality, zoom and pan controls, and color-coded seat types. Supports standing section ticket purchasing, seat preview/mini-map, and performance optimizations including viewport culling and FPS monitoring.
### 3. Advanced Geometry & Flexible Layout System
Curved row geometry with arcs featuring configurable radius and angles, section boundaries with custom paths, and automatic seat naming and labeling. Includes category management for seat types, support for complex venue layouts (stadiums, theaters, event spaces), and background image support for visual context.
These features enable creating complex venue layouts and providing an interactive seat selection experience for customers.

## Geometry Engine Deep Dive
The geometry engine is the core computational layer that transforms abstract layout parameters into precise seat positions. It handles complex mathematical operations to ensure accurate placement across diverse venue shapes, from straight rows to complex curved sections.
### Arc-Based Row Calculations
Arc-based rows require converting angular parameters into Cartesian coordinates. Each curved row is defined by a center point, radius, start angle, and end angle. The engine calculates seat positions by dividing the arc's angular span into equal segments based on the number of seats.
The fundamental calculation uses parametric equations where each seat position is determined by its index along the arc. For a seat at position `i` in a row with `n` seats spanning from angle `θ₁` to `θ₂`, the angle is calculated as:
```
θᵢ = θ₁ + (θ₂ - θ₁) × (i / (n - 1))
```
The Cartesian coordinates are then derived using:
- `x = centerX + radius × cos(θᵢ)`
- `y = centerY + radius × sin(θᵢ)`
This approach ensures uniform angular distribution, which is critical for maintaining visual consistency in curved sections. The engine also handles edge cases like single-seat rows and arcs spanning more than 360 degrees.
### Seat Spacing Along Curves
Maintaining consistent physical spacing along curved paths requires arc length calculations rather than linear distance. The engine computes the total arc length using `L = radius × (θ₂ - θ₁)` and divides it by the desired seat spacing to determine the number of seats.
For non-uniform spacing requirements, the engine supports configurable spacing profiles. Each seat's position is calculated using cumulative arc length, ensuring that the physical distance between adjacent seats matches the specified spacing parameter, accounting for the curvature of the path.
The spacing algorithm also handles variable radius scenarios where rows may have changing curvature. In such cases, the engine uses piecewise arc calculations, maintaining spacing consistency across radius transitions.
### Rotation and Alignment Logic
Each seat must be oriented correctly relative to its position on the curve. The rotation angle for a seat at angle `θ` is calculated as `θ + 90°` to ensure seats face the center of the venue (typically the stage or field).
For straight rows, rotation is straightforward—seats align perpendicular to the row direction. For curved rows, the rotation varies continuously along the arc, with each seat rotated to face the arc's center point.
The alignment system also handles special cases:
- Seats at row boundaries may require adjusted rotation to prevent awkward angles
- Corner seats in L-shaped or irregular sections use weighted rotation based on adjacent row segments
- Standing sections use vertical alignment calculations independent of horizontal rotation
Rotation calculations use normalized angle values (0-360°) internally, with conversion to radians for trigonometric functions. The engine caches rotation values to avoid redundant calculations during rendering.
## Scale & Performance Numbers
Built to handle enterprise-scale venues with thousands of seats while maintaining smooth interactivity.
- Handles 10,000+ seats in a single venue layout
- Maintains 60 FPS during pan/zoom operations on desktop
- Targets 30+ FPS on mobile devices during interactions
- Viewport-based rendering ensures only visible seats are processed
- Memory-efficient strategies prevent browser crashes on large layouts
- Load time improvements of 70%+ compared to previous rendering approaches
This approach allows fine-grained interactivity without sacrificing performance, enabling smooth experiences even with complex stadium layouts.
(Source: https://www.akshatgoel.com/internal/seat-layout-builder)
### Spring Animation (2024)
## Why springs feel good
Hand-tuned cubic-bezier curves move predictably. Springs don't. They reflect the energy of the user's input — a gentle drag returns gently, a flick overshoots and settles. That asymmetry is what makes spring-driven motion feel alive instead of canned.
This component creates a playful, skeuomorphic animation inspired by the work of [Preet Mishra](https://preetmishra.com/craft). Three logo cards drop into a pocket-shaped frame using spring physics — each with its own stiffness and damping so the motion staggers naturally. Hit replay to watch them land again.
## Under the hood
Each card uses framer-motion with `type: "spring"` and a different combination of `stiffness`, `damping`, and `mass`. Higher stiffness pulls harder back to rest; higher damping bleeds energy out faster. Vary either, and the cards land with a different feel — bouncy, floaty, or quick.
The pocket shape itself is just a custom SVG `` applied to a rounded container, so the cards appear to drop *into* the pocket rather than over it.
## A simple playground
If you want to feel the math directly — no bezier, no presets — drag the circle below. Release it and the spring pulls it back to center. Tune the parameters and watch how the character of the motion changes.
## The math
A spring's motion is governed by Hooke's law plus a damping term:
```
F = -k * x - c * v
```
Where `k` is stiffness, `c` is damping, `v` is current velocity, and `x` is current displacement from rest. Each frame, force is divided by mass to get acceleration, then integrated forward in time:
```
v += (F / m) * dt
x += v * dt
```
That's it. No animation curve, no fixed duration. The motion stops when the system reaches equilibrium.
## Why this matters in ui
Drag a sheet down, release, and a spring decides where it goes. Tap a button, and a spring decides how it scales back. Resize a window, and a layout that uses springs continues from whatever velocity the user just gave it — instead of restarting from zero.
This is why ios and most modern apps feel responsive in a way css keyframes don't: the system never lies about where the energy came from.
(Source: https://www.akshatgoel.com/internal/spring-animation)
## FAQ
### Who is Akshat Goel?
A 23-year-old software engineer who came to engineering from motion design. Founding engineer on BNNGPT (auraone.ai), an AI news streaming platform.
### What does Akshat Goel build?
High-performance product tooling (an in-house stadium seat-layout builder), interactive UI and spring-physics motion work, live hardware telemetry, and personal data tools that track running, deep work, and health metrics.
### What does Akshat Goel write about?
Systems and recommendation algorithms — multi-armed bandits, contextual Thompson sampling — and how to evaluate AI agents. The full notebook is at https://www.akshatgoel.com/notes.
### Where can I find Akshat Goel online?
Website https://www.akshatgoel.com/, GitHub https://github.com/akshatgoel07, X https://x.com/akshatgoel0.