Rohan Singh's Weblog

Writing about software, data science, and things I learn along the way.

Bookmarks and the Shape of a Mind

I've been keeping a bookmark library since 2017. It started as a utility — a place to dump links I didn't want to lose. Somewhere along the way it became something else: an accidental autobiography.

I was browsing it recently and noticed something. Across nine years of saving, a shape had emerged. Not the shape I would have described if you'd asked me what I read — something more honest than that.


What I thought I was collecting

Technical references. Documentation. Course materials. Useful tools.

And yes, all of that is there. ISLR in Python from 2019. FastAPI docs. LSF job scheduler commands from when I was learning HPC. DICOM processing tutorials from the early clinical imaging work. Survival analysis notebooks. An annotated PyTorch paper implementation site I still open occasionally.

A large fraction of the library is exactly what it appears to be: a working engineer's toolbox.


What I was actually collecting

Arguments about how the world works.

The Bitter Lesson. Rich Sutton's 2019 essay that almost nobody in industry had fully absorbed even five years later. The core claim: general methods that scale with computation have consistently beaten methods that encode human knowledge. Every time. The implication — that most of what we call "domain expertise" in AI is a liability dressed up as an asset — is still uncomfortable to sit with.

Wigner's paper. The original one. The Unreasonable Effectiveness of Mathematics. I have the PDF. I'm not sure I've ever made it all the way through, but I keep it because I keep thinking about it.

Machine Bias from ProPublica. The 2016 investigation that kicked off the entire algorithmic fairness conversation in a way that academic papers couldn't. The point wasn't just that COMPAS was biased — it was that the bias was hidden inside the thing's own claimed objectivity. That's still the central problem in clinical AI and we mostly pretend it isn't.

The Control Group Is Out Of Control from Slate Star Codex. The best thing ever written about how the scientific process fails while appearing to succeed. Required reading for anyone building AI systems that will touch medicine.


The clinical AI problem, bookmarked in real time

There's a collection called MSKCC with 14 items. Another called Data Science & ML with 123. And scattered across both, a particular cluster: papers and essays about AI in healthcare that are, taken together, a fairly pessimistic literature.

The Shaky Foundations of Foundation Models in Healthcare. Faisal Mahmood's computational pathology lab at Harvard. The Nature Medicine multimodal biomedical AI review. A 2021 paper on the limitations of transformers on clinical classification that still holds.

The throughline: these systems work, and then they don't work where it matters. They generalize in controlled settings and fragment in deployment. The hard problem isn't building the model. It's building the infrastructure around the model that doesn't let it slowly drift into confidently wrong.

I've been inside this problem for years now. The bookmarks are the paper trail.


Feeds I actually read

The Reads - frequently visited collection is the most honest part of the library. No triage, no aspirational reads. Just things I actually go back to.

Danluu — Dan Luu writes software engineering essays that are empirically dense and take nothing on faith. Every claim cited. Every intuition tested. The longest posts are the best ones.

antirez — Salvatore Sanfilippo, who built Redis, now writes about programming, life, and the strange experience of having built something that outlasted the version of you that built it. One of the few technical writers who sounds like a person.

Stratechery — Ben Thompson on the business and strategy of technology. I disagree with him more than I used to. I still read everything.

Simon Willison — The most generous technical blogger on the internet. Writes about LLM tooling with the energy of someone who finds it genuinely delightful. Useful for staying calibrated on what's actually possible vs. what's being sold.

Daniel Lemire — Computer science fundamentals, performance, algorithms. The posts that make you realize you don't understand something you thought you understood.

Drew Breunig — AI commentary without the hype. Quietly one of the best analysts writing about where this is actually going.


The breadth problem

Here's what the library reveals that I wouldn't have said about myself: I am constitutionally incapable of staying in one domain.

Machine learning textbooks sit next to the PageRank paper. Fairness in ML next to cognitive bias lists. An R forecasting book next to solo travel guides for Paris from 2014. The Reasons to Be Cheerful blog. McSweeney's Internet Tendency. A cartoon about a moving mind. The FiveThirtyEight Riddler puzzles. Retraction Watch.

This is either a feature or a flaw depending on the day. On good days, unexpected connections show up between things that shouldn't connect. On bad days, it just means the Data Science & ML collection has 123 items and zero are tagged after 2023.


A note on timing

The oldest bookmarks are from when I was a student. The CMU course videos. The intro stats pages. The learning-resources aggregators. That version of the library was about becoming — collecting the building blocks of a person trying to get somewhere.

The middle period is all tools and production systems. Kafka ML integration. MLOps pipelines. Label Studio. QuickUMLS. Things you save because you need them tomorrow.

The recent stuff is different. More essays. More blogs. More things that don't obviously connect to any particular deliverable. The Boil the Ocean piece I saved yesterday — about the failure mode of strategic overcomplication — is not going to help me ship anything. I saved it anyway.

I think that's the shift. Early career you collect tools. Later you collect perspectives.

The bookmark library is a record of which version of you was running at any given time.


What I should read next

Of everything I've saved and not read, a few things are aging in the right direction:

  • Dario Amodei's Machines of Loving Grace — sitting in Read Later since January 2025. His argument about AI and medicine specifically. I've started it twice.
  • Nathan Lambert's RLHF Book — the most comprehensive treatment of post-training that exists right now. Saved in April 2025. Directly relevant to work I'm doing. No excuse.
  • Soumith's "How to train on 10k H100s" — practical infrastructure wisdom from the person who built PyTorch. Saved in October 2024. Still relevant.
  • The JAX scaling book — free, comprehensive, and I've been skimming it when I should be reading it.

The library has 363 active items and 343 in the trash, which is its own kind of poetry.

Most things you save, you don't come back to. That's fine. The act of saving is itself a small declaration: this seemed worth keeping. The aggregate of those declarations is something like a portrait.

I'm not sure I would have predicted this portrait if you'd asked me at the start. But looking at it now — the fairness papers next to the infrastructure docs next to the essays about why science breaks next to the Redis creator's blog — it feels accurate.

The shape was there. I just hadn't looked at it all at once.


If you're building anything in clinical AI and want to compare notes, I'm around.

← Back to writing