Tabletop Library

NEWS

Updates, announcements, and stories from Tabletop Library

Nabeel Hyatt

Making a Better Board Game Recommender

It started harmless enough, with asking ChatGPT what board game to play. Then I decided to spend a weekend to improve the results a bit. Seven months later, we're emerging from a very deep rabbit hole on what it means to recommend something.

TLDR: We just launched a board game recommender, you can try it here. You tell it who is playing and what you’re in the mood for, and it’ll suggest twelve games from our shelves, along with reasoning about why those are good picks for you and your group.

A few years ago, someone like me could never have built this. And more likely, a place like Tabletop Library would never even try. These recommenders are generally the domain of a team of PHDs at Netflix and Spotify with a billion data points. We, on the other hand, are a local board game club that hasn't even opened yet, with a few hundred members, a few thousand games, and some weekend free time.

So consider this a peek behind the curtain at what building software looks like right now, from an amateur feeling his way through a field real experts spend their careers on. (If you are one of those experts, I'd love your feedback)

It all started with a reasonable question to ask an AI in our current age:

  • Four of us want a visually beautiful game driven by cards that we can finish in under 60 minutes. What should we play?

The first time I did this it kind of worked. It's was even a little magical. But I found that like a lot of AI stuff, repeated usage really shows the difference between something that quickly kinda works, and a really useful tool.

If you ask this same question a couple dozen times, as I did late last year of Claude and ChatGPT, you notice some things:

  • Claude recommends Wingspan virtually every time it was asked. This makes sense. Wingspan is beautiful, full of cards, and one of the most acclaimed board games of the last decade. Claude said it would take 40 minutes. Stats from BGG say 70 mins, my experience says it's more like 90 if you are teaching it.
  • It often recommended Azul, a game with no meaningful card play, because it offered a “card-adjacent drafting experience.” This is an excellent phrase to write when you know Azul is a lovely game and would prefer not to dwell on the cards question.
  • One run included Lost Ruins of Arnak, listed in the same answer at 120 minutes. The request said 60.
  • It also recommended Jaipur twice. And yes it is full of cards, though I'm not sure I'd call it beautiful. But the kicker for this four player recommendation was this description, "a great 2-player classic, so it’s not ideal for all four of you at once, but it’s a strong “beautiful and fast” option if you split into pairs."

I'm not cherry picking failures, they are indicative of the trend. But still, if you typed that into Chatgpt or Perplexity or Claude you could easily read the list, pick Wingspan, and move on with your life.

However the answer still left out a few things for us at Tabletop Library:

- Do we own the game?
- Is it actually best with four people, or just kinda playable at four people?
- Does anyone in the group know how to teach it?
- Does one person in our group hate this type of game?
- Is “under 60 minutes” a preference, or does someone need to catch a train?
- Are we recommending Wingspan because it is right for this group, or because the internet has written a lot of sentences about Wingspan?

That last question is more complicated than it looks. Research on zero-shot LLM rankers has found sensitivity to both item popularity and to things as subtle as where a name appears in a prompt.

To research further I gave Claude Sonnet and GPT twelve prompts 50 times each (stuff like "games where we negotiate a lot" and "best co-op games for beginners and advanced players together). Across 600 recommendations, it returned only 61 unique titles. Blockbuster popular games like Azul and Pandemic appeared constantly.

Repeated prompts overlapped by 58% on average. For the opening question listed here, the overlap was 37%. That means answers were both repetitive and unstable, which is an efficient way to acquire two problems at once.

So I decided it could be fun little side project to try and build a better board game recommender. After all, it would be a nice way to test out some frontier LLM stuff, maybe do some custom model building, explore vector databases, and so on for a few weeks.

I mean, how hard could it be?

Oh boy.

The first recommender was basically an AI agent

I should own the ridiculousness of this up front. I spent seven months deep-diving a subject I don’t know much about, to build an esoteric tool whose best-case outcome is a few hundred board game club fans being mildly happier on a Tuesday night. It is perhaps a good example of of what can happen with vibe coding though. It's not so much that this is replacing what in the past a software engineer would have built, it's that this just would never have been built at all.

I drafted the first version of our `/play` recommender in December 2025, and it was relatively typical for what you would imagine of the time.

First, an LLM (GPT 4-o back then) was given the search term, and separated things that sounded like firm constraints ("four players") from things that sounded subjective (like “beautiful” or “cozy”)

Then I built a vector database of all of Tabletop Library’s catalog with TurboPuffer, using a mix of semantic and keyword search. It fetched up to 30 plausible games from the database with player counts, durations, complexity, descriptions, our classification data, and anything we knew about the players.

Then we handed the whole packet to GPT-4o again and said, more or less: pick the best 12 based on the prompt you were handed.

It also wrote a reason for every recommendation. Stuff like: "David can teach this one." "Molly has been wanting to learn that one." "This game has the same auction mechanic as a game you liked."

This was a sensible architecture. It was also essentially the first architecture many people build for an AI agent:

1. Retrieve the available context.
2. Put it in a prompt.
3. Ask a capable model to figure out the rest.

And the results were familiar to anyone with a pattern of working with LLMs. As a first shot, it worked better than it had any right to, which felt great. But the gap between "okay as a first attempt" and "actually great" did not seem to be improving at all as I was adjusting LLM prompts over and over.

The issue was the model still owned nearly every important decision after retrieval: constraint adherence, ranking, group compromise, novelty, and the final explanation. One opaque LLM call decided whether a 70-minute game was close enough to a 60-minute request and whether one player’s enthusiasm outweighed another’s dislike.

From here I decided to do a little research. Recommendation systems are not new after all, so this has got to be a solved problem we just need to find the right approach?

Board games are a tricky little recommendation problem

There are three core issues that made most recommender systems that you see in the world not a great fit for this little cardboard hobby.

1) Most of the best recommender systems were built for abundance.

Music services have tens of millions of tracks and enormous listening histories. Netflix built its recommender around a vast catalog and an even larger stream of views, searches, ratings, and browsing behavior. There's a long line of academic literature, and the vast majority of the innovations in machine learning on recommenders are leaning into the amount of data available.

But there are relatively fewer board games in the world than songs. Even more acute, Tabletop Library has a couple thousand games and a membership measured in hundreds. Many members have rated fewer than ten games. Some have rated none. A new member may know that they enjoy “card games” or that Catan was pretty fun eight years ago. But that is not enough data to train a small regional Netflix.

2) Board games are also consumed differently.

If Spotify recommends a bad song, you skip it. Four minutes have passed, assuming you were very patient. If a board game recommender makes a bad choice, four people may spend 20 minutes learning rules, another 90 minutes playing, and the drive home discussing why nobody stopped David from choosing it.

3) Then there is the group itself.

Most recommenders answer a personal question: what will this person click, watch, or buy? But a game-night recommender answers a temporary social question: what will these particular people enjoy together, tonight?

The last thing you want here is the highest average rating. If three players are delighted and one player will be miserable, the averaging rating may still look respectable. But faces of players at the table will not be.

Other people have been working on this for a while

So I started digging into the academic literature for stuff we could use.

Recommenders began as far back as Xerox PARC’s Tapestry, but the first time they really entered my awareness was the early Internet era with music services.

In these pre-Spotify days, the three early major services were Last.Fm, Pandora, and Yahoo! Music, and each took a different approach to recommendation that still loosely defines the paths used today.

  • Pandora invested in knowing the content, basically data as the key. Its Music Genome Project employed trained musicologists to describe hundreds of details about songs like the tone, instruments used, etc.
  • Last.fm invested in social behavior: its “scrobbling” system recorded what people actually listened to, producing taste histories that could be compared with other listeners.
  • Yahoo! Music combined explicit ratings with a hierarchy spanning tracks, albums, artists, and genres. The Yahoo! Music dataset used in the 2011 KDD Cup contained hundreds of millions of ratings from more than a million anonymized users.

As an avid collector of MP3s in those days and I remember being fascinated by the inherent contrasts in style. Content-based systems can understand a new item before anyone has listened to it, but they need good descriptions. Collaborative systems can discover relationships nobody thought to label, but they need enough overlapping behavior. Explicit ratings are clear, but most people do not want a second job rating everything they encounter.

These three basic architectures actually still define the major approaches today. But now they are implemented as hybrid systems. Robin Burke’s survey of hybrid recommenders described several ways to combine systems so that one could cover another’s weaknesses. Netflix eventually made the same point at production scale: its recommender system was not one algorithm. It was a collection of ranking systems, rows, search behavior, contextual signals, offline tests, and online experiments.

So I started weighing the various recommender systems in the academic literature against each other and applying them to our problem.

For example, the earlier issue of not using the average happiness rating of all the players because one person might be miserable for 90 minutes? That's called least misery scoring.

There are other approaches: average preference, multiplicative scoring, fairness rules, taking turns serving different people. Judith Masthoff’s work on group recommenders found that people use several of these strategies.

So many papers! And there's no specific formula handed down on a stone tablet for our problem. However, this is the kind of complex, academic research that is well suiting for Claude to spend a day doing deep research on.

Asking Claude "hey what's the academic literature on this subject" ended up yielding pretty surface level answers without much insight. But as things got more agentic, it became possible to ask Claude Code (in Conductor) to spin up a bunch of sub agents to read every single article and paper on Recommender systems, then write a text summary of each article analyzing how it could potentially help our particular problem.

That, combined with using academic AI research tools like Elicit, yielded directories full of architectural comparisons: the systems we could build, judged against the published literature for our exact problem: a couple thousand games, a few hundred members, sparse ratings, group play.

Armed with this, it turned into a wonderful enjoyable rabbit hole just before bed each night reading, researching, and rapid prototyping.

The five stages of the recommenders

Five different recommender approaches made the cut, systems that combined the various recommender styles: some leaned on LLMs like ChatGPT and Claude, some on deterministic scoring formulas out of the search-engine literature, some on the old social, people-like-you recommenders. The permutations did get quite complex, but in the end we ended up with a relatively simple approach that beat them all.

The five, in order:

(1) The Simple LLM Ranker

You have already met the Simple LLM Ranker. Retrieve candidates from our database, hand everything to the AI model. This is the first system we built. The story of the other four is which responsibility moved out of the model at each step.

(2) The Multi-Source Ranker and Retriever

Let basic code decide what games are in the pool to decide from

The first major change — and the one that survived into every version after — was a deterministic hard-filter gate. Even if we were going to use an LLM, it needed to be given a very clear sandbox to play in.

If someone asks for a four-player game, we should write code to make sure the game actually is 1-4 players. The same applies to complexity, excluded games, required mechanics, explicit categories, stock, and other constraints.

Using multiple approaches to generate potential games

The original system also got candidates from one place: search. This meant the ranker could only choose games the search system happened to retrieve.

We added three more channels:

- games on the seated players’ lists;
- games liked by members with similar rating histories;
- games related to a title named in the request.

Two of these channels are the music-service ideas at club scale. “Members with similar rating histories” is Last.fm’s move: find people whose history overlaps yours and borrow their enthusiasm. “Games related to a title” is the item-based variant of the same trick: two games are similar if the same people like them both.

These methods are not there to rank anything, it just decides which games are allowed into the conversation.

(3) The Deterministic Scorer

Then we tried a recommender that is more akin to what you would have found prior to this AI craze, with no LLMs at all. We built a Deterministic Scorer with a weighted score. Retrieval rank, collaborative evidence, list affinity, semantic fit, quality, novelty, group support, and duration fit each contributed a visible amount.

The recipe came from search engines, where it is called reciprocal rank fusion: trust no single list; a game earns points for ranking high on several lists at once.

This was fast, reproducible, and inspectable. The downside it that is it a bunch of inscrutable weights - something is a .12 on novelty and .10 on player match and just moving them around all day felt fruitless.

If we had hundreds of thousands of perfectly taste aligned rankings, then we could automatically tune these values and maybe it could help. But remember, we've got a sparse data problem here. We don't have the kind of users that Netflix has.

Additionally, hand-tuned math weights doesn't really help solve subjective board game requests like “something pretty where I use cards.”

(4) The Specialized Reranker

I had heard about a startup called ZeroEntropy that was supposed to have an excellent reranker, and it seemed like it could perhaps help with the problem we were having with the Deterministic Scorer. It's a model built specifically to rank documents against a query. A reranker does exactly one thing: given a request and a description, it produces a match score. The goal here was to combine the Deterministic Ranker from the prior step, while still allowing some natural language like "i don't want a mean game."

It cannot chat, cannot explain itself, and does not know what a board game is, but it's way cheaper and more tunable than an LLM like Claude. It was a good system, however, it was not the best system in our comparison.

This is an important category of engineering result, in that it was fine, but I just found better things through iteration. Although we did move from this as the primary judge, we found this approach was quite good at doing group support so we use do use the signal still.

But in the end, we got better results just handing that signal to an LLM to make final calls.

(5) The Librarian: A Hard-Gated, Taste-filled, LLM Judge

The winner started from a different question. The day Anthropic's Fable came out everyone was talking about giving it big nebulous problems because the new model was great at exploring problems. So, I gave it access to all the academic research we had done and asked it a simple question, "why do AI models have bad taste when it comes to this problem?"

It didn't exactly spit out the answer, but from that discussion it was clear that maybe the problem wasn't lack of taste or judgement, but that we lack the vocabulary to describe something as nuanced as playing a game.

So instead of building a large scoring formula around every available signal, the challenge pivoted to figuring out how to generate better data, and handing it to the LLM in the right format. That took two forms:

  1. A structured taste profile for each member - think of it like a private player dosier that says "Jamie likes games that teach quickly, their favorite games are Jaipur and Mahjong." and so on.
  2. A private dosier of each game describing not the ratings, or hard constraints like number of players, but how the game feels to play. How luck based is it? Is it highly interactive with other players or like playing solitaire?

This handled our actual data sparseness much better. It also let us encode negative evidence. A one-star “Not For Me” rating is an absolute veto. Strong, well-supported opposition is so important for us: if a member’s history says they consistently dislike direct conflict, the table does not get a knife-fight negotiation game no matter how enthusiastically the other three would vote for it.

The group score then uses a geometric mean across members with confident profiles. This is related to the multiplicative strategies in the group-recommendation literature: one low fit pulls the result down more strongly than it would in a simple average. Three members at 0.9 and one at 0.2 average to a respectable 0.73; the geometric mean says 0.62, which is the arithmetic version of glancing at the fourth player’s face. We retain explicit vetoes for cases where “pulls the score down” is not strong enough.

So the Judge is misery-aware, but it is not literally least misery.

For our problem (a couple thousand games, a few hundred members, sparse ratings, and groups instead of individuals) I think a hard-gated LLM judge is the right shape: code owns feasibility and group math, a hand-built taxonomy supplies signal the internet does not have, and the LLM judges a bounded slate it cannot smuggle anything into.

It is still a system just a beta, if you will, but the results have been very promising and the best we've tested.

What the current system lets the LLM decide

The "Librarian" (a Hard-Gated, Taste Filled LLM Judge) is still an LLM system. We just gave the LLM a narrower job.

Here is what happens when someone makes a request.

1. Separate hard constraints from subjective mood

“Four players” and “under 60 minutes” become structured constraints. “Beautiful” and “driven by cards” become semantic intent. If the member uses the visible filters, those values are authoritative. The query-understanding model can interpret natural language, but deterministic parsing handles recognizable phrases and provides a fallback.

There is a real difference between “I generally like shorter games” and “the babysitter leaves at nine.” The database cannot infer the babysitter, but it can stop treating every number as a mood.

2. Gather game ideas from multiple sources

TurboPuffer searches the catalog. Player lists contribute games someone can teach, wants to learn, or enjoyed before. Collaborative filtering contributes games liked by people with similar ratings, when enough rating data exists. Reference expansion contributes games related to a title in the request.

3. Throw out games that don't match

The merged pool passes through deterministic constraint checks to form a "perfect" pool of games to pick from. If that pool is too thin, say it generates less than 40 games, then we have a "near-match" policy that can add a small number of close games that were close but not quite.

The LLM does not get to declare Lost Ruins of Arnak a 60-minute game through force of personality.

4. Compute fit for each person individually

Every member profile contains a distribution across our game classification, preferences across five gameplay dimensions, complexity and duration ranges, negative signals, and a confidence score.

For each candidate, code calculates how well it fits each member. It applies list and rating evidence, detects vetoes, finds teacher/learner pairs, and checks whether a difficult game has someone capable of teaching it. Stationfall is a great recommendation at a table where someone knows it, and a hostage situation at a table where nobody does.

Then it combines the per-member scores into a group score.

5. Reason on what people asked for

Stable taste is not the same as current intent. Someone who generally enjoys three-hour spreadsheet simulator may currently want “something light because I am tired.”

This was obvious to see after the Hard-Gated LLM Judge produced candidate pools that barely changed when you typed in different text. The profile signal was doing its job too enthusiastically.

Now, when the request includes text, the system blends group taste with query relevance before selecting the games for the LLM to judge. If you leave out mood text, it remains purely profile-led.

6. Ask Claude to rank a specific slate of games

Claude receives at most 40 eligible candidates, the member profiles, per-member fit, the request, and supporting evidence. It returns candidate indices rather than repeating long database IDs, ranks up to 12 games, and writes one short reason for each.

This is where LLMs are useful. They can compare several kinds of soft evidence at once and understand that “pretty cards” may refer to illustration, physical production, tableau building, or simply wanting to hold a hand of cards. They can also turn structured signals into a sentence a person might want to read.

7. Check its own work

Code then maps each returned index back to a real candidate. Invalid indices disappear. Duplicates disappear. Anything that failed the deterministic gate disappears. If the model returns too few games, deterministic ordering fills the remaining slots.

We also scan the reasons for internal vocabulary and incorrect product language. If Claude tells a member about their “aggregate fit,” “z-score,” or `teacher_gate`, that game gets deterministic copy instead. After all, members came to choose a game, not attend the recommender’s sprint review.

Finally, the system sorts perfect matches ahead of close matches again. Defense in depth is what you call doing the same sensible thing in several places because you have met software.

The useful data is the data we made ourselves

The largest general model does not know Tabletop Library’s collection as well as Tabletop Library does. More importantly, it lacks the vocabulary we need.

BoardGameGeek has an enormous amount of useful data: categories, mechanisms, ratings, complexity, player ranges, and descriptions. But categories such as “strategy” and “family” are broad. “Card game” tells you what components are present, but not what playing feels like.

BGG has also wrestled with sparse ratings, and its fix is worth knowing. Bayesian averaging is a popular technique—BGG’s Geek Rating uses it—because when a game has only a handful of votes, one person’s subjective opinion swings everything. So BGG seeds every game with about 1,500 imaginary ratings of 5.5, and real votes have to beat the ghosts. It is a good trick. It also felt like it didn’t get to the root of the issue: smoothing tells you how good a game is on average, not whether it is right for these four people tonight.

So we started to wrestle how to describe a games taste and feel. Thankfully, we'd already been working on Tabletop Library Classification System, or TLCS, to organize games by their primary form. It is our board game version of a library call-number system. A drafting game, a campaign game, a social deduction game, and a heavy negotiation game live in meaningfully different neighborhoods.

We also classify every game across five gameplay dimensions:

- Thematic depth: abstract, themed, or immersive
- Randomness: luck-driven, tactical, or skill-driven
- Player interaction: minimal, indirect, or direct
- Learning: intuitive, moderate, or heavy
- Tempo: light, thoughtful, or intense

These describe play rather than mechanics or marketing. A member may not know the phrase “multiplayer solitaire,” but they know they dislike spending 90 minutes building a private collection of cards while three other people do the same thing nearby never interacting.

So our member taste profiles now could turn each member’s ratings and lists into distributions across this vocabulary. A five-star rating or “can teach” signal carries more weight than merely having played something. Low ratings become negative evidence. The profile also records how much evidence it has.

This is our most important advantage. It is not really the prompt or even a model. It is a small, opinionated body of domain knowledge that does not exist anywhere else in quite the same form.

What is still wrong

The new recommender is better. But, it's also built on several suspicious facts.

  • Balance between games we know you love, and new games. Yes, you may really love playing Mahjong, but that does not mean you want Mahjong recommended every single time. In fact, you might be in the mood for a new game. Right now we're still tweaking with how repetitive vs novel it is, how much it should use the Play List of games you are interested in vs find a discovery pick.
  • How long does a game really take? Our catalog has a duration for each game, usually derived from publisher or BoardGameGeek data. This is useful but plainly very inaccurate. Eventually, duration should be a distribution conditioned on player count, how familiar the group is with the game, and whether someone can teach—not one integer printed on a box.
  • What's the right player count? 1-4 players on the box can often mean playable at 1, but amazing at 3 players. BGG already collects community votes for whether a game is Best, Recommended, or Not Recommended at each player count. We'd love to integrate that into the rankings so you finding the best 3 player game for your 3 player game night.
  • Did the recommendation work? Our evaluation mostly measures whether a recommendation looks relevant. The real outcome happens at the table when the game is picked. We have no doubt once we can actually connect this to real members picking the games, and then how they rate them afterwards, we can get better quickly.
  • Does the list contain twelve versions of the same idea? Relevance tends to cluster. Ask for a medium-weight nature game and a recommender can return a tasteful wall of green boxes. Information retrieval has long used methods such as maximal marginal relevance to trade a small amount of pure relevance for diversity, and there is a related line of work on re-ranking away popularity bias. We do not yet apply either to the current system. We probably should. The nature of a library is a little bit of exploration, after all.

Did we make a recommender harness?

It would be satisfying to end by saying we replaced the LLM with good old-fashioned software engineering. We did not.

If we were trying to raise seed funding from our friends I'd probably call what we ended up with a harness. But since we are just a local board game store looking to break even, I'll put it in plainer language. The current system very much uses an LLM to understand loose requests, compare nuanced candidates, and write useful reasons. LLMs are just much better at weighing the kind of subjective choices inherent in board games than older ML recommendation systems.

What changed was the boundaries it operates in.

The model does not decide whether four is between one and three. It does not get to override a member’s “Not For Me” rating, or silently put close game matches above perfect ones. It chooses among candidates that we've decided it is great to choose from, using domain knowledge that the general internet does not contain, and then code checks the answer.

This was also the lesson for building software with AI agents that we'll likely take forward in other systems. Chatgpt can make a first recommender surprisingly quickly. It can wire up vector search, write the prompt, return a nice grid of cards, and add a tasteful loading state while it is there.

That is real progress. It is just not the same as knowing what a recommendation should optimize, and it's definitely not the same as knowing the difference between good and great.

The longer work was deciding which facts mattered, creating a vocabulary for taste, building comparisons that could themselves fail, and moving each decision to the part of the system best able to make it. I'm sure there's a lot more to evolve in that vocabulary as we keep building.

A system can only judge what it can name, so before anything (model or code) could weigh how an evening would feel, someone had to sit down and encode feel. From there it was great to see Claude had excellent taste, it just needed the language to understand what we want.

Which brings me back to the ridiculousness.

People have always poured unreasonable hours into esoteric things that make a small number of people happy. Every zine, every hand-painted miniature army, every homebrew RPG campaign lovingly written for an audience of six.

What’s wonderful is that custom software can be one of those things now. A recommender of this size and scale went from something a large corporation would fund, to something a local membership club could build just for the joy of it.

That joy still did involve a ton of hours making judgement calls and iterating -- despite the online hype about one-shotting everything with AI that was certainly not the experience here. But still, it is well within the scope of a hobby project all the same.

And if you build recommenders for a living and have been wincing since paragraph three: good, you’re the other reason I wrote this. Tell us what we got wrong. I’d love nothing more than an email that starts “have you tried…” as I suspect we'll have plenty more fun building this from here.


--

PS: If you want a four player, beautiful card game that plays in under 60 minutes, give Glow a shot.


---

Sources and further reading

What we read in May 2026, before most of the rebuild: Burke (2002) on hybrid recommenders; Sarwar et al. (2001) on item-based collaborative filtering; Cormack, Clarke & Büttcher (2009) on reciprocal rank fusion; Carbonell & Goldstein (1998) on maximal marginal relevance; Masthoff on group recommendation strategies; and Abdollahpouri et al. (2019) on popularity-bias re-ranking.

Plus the industry precedents: Spotify’s Discover Weekly (a hybrid explicitly to cover cold start), Goodreads’ ~20-ratings gate before collaborative filtering turns on, and BGG’s Bayesian Geek Rating.

For the longer history: Goldberg et al. (1992) on Tapestry; Resnick et al. (1994) on GroupLens; Shardanand & Maes (1995) on Ringo and “automating word of mouth”; Gomez-Uribe & Hunt (2015) on Netflix; Herlocker, Konstan & Riedl (2000) and Sinha & Swearingen (2002) on explanations and transparency; Hou et al. (2024) and Lichtenberg et al. (2024) on LLMs as rankers.

Vera Devera

Opening Day: We're Getting Close!

Opening Day: We’re getting close!

Last Wednesday — one year, four months, and nineteen days after filing (but who’s counting?) — our permit was approved and construction is underway! Thank you to the City of Berkeley for taking on this massively risky bet and approving a business where people will sit in chairs and play with cardboard.

There’s one more Final Boss Permit after construction is complete, but if all goes well (and why wouldn’t it?), we will be open by August.


Get a Founding Membership before we stop selling them on Friday

Founding members (now over 100 strong!) get special perks: Your name on a plaque, an exclusive t-shirt & bag, and we won’t change your membership price for two years. For context, while we’re trying not to muck with prices until we’re open (we want to see how frequently people come, and thus, how many memberships we can offer), but we will almost certainly muck with prices.

If you care about those things, you can still be a founding member, just join by this Friday, May 8th.

Become a Founding Member


Launched Today: TTL @ Home - Recruit TTL members to your game nights

For board game lovers, it can be a challenge to find people that want to play the same game as you… not everyone has a family eager to play a four hour session of “Sheep Shearing: 1427.” Tabletop Library wants to help solve this by offering tools that make it easy to find and organize games with other members.

Today we’re launching TTL @ Home. While we await our official opening, members can use our app to organize a game night somewhere else. By using TTL @ Home, you can:

  • Recruit other TTL members to your game night
  • Borrow a game from our library
  • Use our “whenisgood” style interface for coordinating a time with other players

Here’s a short video walkthrough

Try it out and let us know what you think. It’s new and may be rough around the edges, but if you approach it as an app built by a corner ice cream store — not a tech company — you’ll be blown away.


Nabeel Hyatt

TLCS: How to Find a Game When You Don't Know What You're Looking For

We have 700+ games now in our collection ready for day 1. That's the good news and the bad news.

Choice is an opportunity and a problem. If you walk into the average board game store having played Catan and maybe Ticket to Ride, you look up at the wall of boxes and you have no idea where to start. Everything has a dragon or a spaceship on it. The names mean nothing. Someone behind the counter asks if you need help and you don't even know what question to ask. You end up buying an expansion for something you already own and leaving.

Most game stores organize alphabetically, which doesn’t help, or with a few broad categories like “Family” “Strategy” and “Party.” These are a decent start but these group a game like Azul, which takes 30 mins and anyone can learn, with a game like Twilight Imperium, which takes eight hours and can end friendships.

Online people might use a search engine, but we realized for physical browsing libraries had figured out a great method centuries ago. They created systems where you could wander into a section, find something interesting, and realize there's a whole category of things you never knew existed. So we borrowed that idea.

The Tabletop Library Classification System

TLCS. A system to make sense of the expansion world of gaming. Every game in our library gets a number. Ticket to Ride is 470.2. Each part of that number tells you something different about the game. The genre, the mechanics, the complexity, and how it actually feels to sit down and play it.

The Categories: What Kind of Game

The first digit tells you the broad genre.

But these aren't organized by theme, you won't find a "Fantasy" section or a "Sci-Fi" shelf. Instead, the categories map to the question you're actually trying to answer: What kind of experience am I in the mood for tonight?

Want something where you're all on the same team? That's the 300s. Want to quietly optimize your own little engine while your friend does the same? 400s. Want direct conflict where every move might ruin someone's plans? 500s.

You've already narrowed 800 games to 150.

100s – Classic & All Ages (Chess, Crokinole, Azul) Something everyone at the table can play. Your parents, your kids, your friend who "doesn't like board games." Simple to learn, still satisfying to master.

200s – Social & Party (Codenames, Mafia, Wavelength) Games for groups. Hidden roles and accusations. Generally, things that are more fun the sillier it gets.

300s – Cooperative (Pandemic, Spirit Island, The Crew) Everyone wins or everyone loses. You're battling the game together, not each other.

400sEuro Strategy (Wingspan, Agricola, Ticket to Ride) Build your own thing without someone knocking it over. Optimization, resource management, satisfying combos. You're competing, but mostly by being better—not by attacking.

500s – Competitive (Munchkin, Root, Twilight Imperium) Direct conflict, tough choices. Territory control. War games. Every action could hurt your opponent as much as it helps you win.

600sNarrative & RPGs (D&D, Gloomhaven, Sleeping Gods). Experiences over mechanical optimization. Get lost in a story. Campaign games that remember what happened last session. Characters that grow.

900s – Special Collection

Puzzles. Prototypes. Teaching copies.

(The 700s and 800s are currently empty. Room to grow.)

The Subcategories: How it Works

The second digit tells you how it works—the core mechanic, the thing you're actually doing over and over.

Ticket to Ride is in section 470, which is called Networks & Routes. Every game on that shelf shares the same basic engine: you're building connections between places, claiming routes, sometimes blocking paths.

This is the real unlock. A new player doesn't walk in asking for "a network game"—they don't know that's a thing. But after they play Ticket to Ride and love it, they can find it on the shelf and look around. Airlines Europe. Brass: Birmingham. Power Grid. Age of Steam. Same family, different complexity levels and styles.

They've just discovered a genre they didn't know existed. Same way you might realize you're really into noir films after stumbling into one.

Complexity: How Heavy Is It

Then every game gets a decimal from .1 to .5. This tells you how much it's going to ask of you.

  • .1 – Light. Codenames, Uno. Teach it in two minutes, play it in fifteen. Your non-gaming friends will be fine.
  • .2 – Medium-light. Scrabble, Sky Team, Azul. A little more to chew on, but you can still learn as you go.
  • .3 – Medium. Mahjong, Wingspan, Chess. Real decisions, real depth, but won't melt your brain.
  • .4 – Medium-heavy. Viticulture, Agricola, Brass. Set aside time with the rulebook, or get a teach. But can really be worth it.
  • .5 – Heavy. Lisboa, Spirit Island, Twilight Imperium. Bring snacks. Bring patience. This is an event of interlocking rules and mechanics to test you.

The nice thing about this: say a friend recommends Lisboa (460.5), which is great but also a .5—a genuine commitment. You can walk over to the 460: Worker Placement section, find Lisboa, and then look left. There's Everdell at .3, Stone Age at .3, Fabled Fruit at .2. Same family of games, easier entry points. Start there, work your way up. It's a path, not a cliff.

The Gameplay Taxonomy: How it Feels

Those four numbers tell you what category a game belongs to and how complex it is. But some games in the same category can still feel completely different to play.

So we added one final layer—five dimensions that describe the how the game feels to play. You can find these by looking up the game on our website, and they describe five essential axis to how a game feels.

Theme: Abstract → Themed → Immersive

Is the game purely abstract like backgammon, or is there a world to get lost in?

Randomness: Luck-ish → Tactical → Skill

Is playing about rolling dice and hoping, or is every outcome a result of your decisions?

Interaction: Minimal → Indirect → Direct

Are you basically playing solitaire next to someone, or stealing their stuff and ruining their plans?

Learning: Intuitive → Moderate → Heavy

Can you pick this up as you go, or do you need someone to teach you first?

Tempo: Light → Thoughtful → Intense

Quick and breezy turns, or deep tension between every move?

These are meant to add texture to the classification. Two games might both be worker placement, but if you're tired late in the evening and someone hands you something marked Tempo: Intense, you might want to steer them toward something else.

Where this is headed

Here's the scenario we're building toward. First, even if you don’t know a single game, you can start in a section based on how you want to play and feel confident in a good choice. Second, once you’ve found a game you like, you can explore that subcategory and maybe find a new favorite.

And lastly, our long term goal is some future date where we have a real sense of the kind of games you enjoy. We would love to get to a point where you can come in with a friend, a spouse, and maybe another member you've never met. Four people, four experience levels, probably a dozen different preferences.

If we know what games you've each played and what you liked about them, we can look at your profiles and say: here are three games none of you have tried that you'll probably all enjoy.

That's the difference between "we have 800 games" and "here's a game for you."

Vera Devera

Opening Update & Fractured Sky: Awakening Preview with IV Studios

Opening Update

I hope 2026 has been off to a great start for you. Unfortunately, we’ve run into more unexpected permitting roadblocks. Our staff and construction team are standing by, ready to begin work as soon as we get the go ahead from the City of Berkeley, but sadly this means we’re unlikely to open before late spring. I'm grateful for your patience and will keep you informed as we move forward!

Cheers,

Vera


Upcoming Event: Fractured Sky Awakening Preview with IV Studios

We’re excited to announce that Tabletop Library is partnering with IV Studios for an exclusive demo of Fractured Sky: Awakening on Thursday, March 12, from 5 to 9pm. The event will include a how-to-play tutorial for the base game and its expansions–plus Soar, a small box game launching with the campaign. RSVP 


Membership Spotlight

Recently, NPR aired a segment on why simply saying “we should hang out” won’t lead to real friendships. It got me thinking about why tabletop gaming is such a fantastic way to meet new people: it reduces the social pressure of always being “on” because you can focus on the board in front of you—and yet, you're sharing the highs and lows of an epic (or silly) game over the course of an evening.

That segues into how excited I am about our members. The Tabletop Library community is shaping up to be as diverse as the games on our shelves! Whether you're a journalism junkie turned aspiring yachtsman, a book club organizer venturing into wooden token crafting, or someone who treasures the puzzle of finding optimal plays—you'll find your people here. We’re excited to be your third space and connect you with new folks to game with. 

Join as a Founding Member today 


News From the Collection

This month we’re highlighting some of our favorite two-player games that are easy to set-up and play, making them perfect after a hectic day at work, or over a relaxing weekend:

  • Fromage: A game where you craft artisanal cheese–simultaneously. The board rotates between rounds, revealing new opportunities to score, and forcing you to adapt as different sections come into focus.
  • Patchwork: A cozy puzzle classic, now sporting a refreshed, modern look. On your turn, draft and stitch tiles together; the most efficient quilt earns the most points!
  • White Castle Duel: A two-player version of The White Castle, a strategy game set in feudal Japan. Place courtiers, and manage resources and actions in an attempt to gain influence  and shift the balance of power.  
  • Iliad: In this head-to-head game of cat and mouse, you will place tiles in alternating patterns to gain the favor of the gods.

Pre-orders: We’re looking forward to stocking our collection with some highly anticipated games such as The Old King’s CrownOrlojAgent AvenueRecall, and more. If you’re interested in pre-ordering with us, reply to this email with the games you’re interested in (and members get 10% off)!

Vera Devera

Construction progress and gifting ideas

As we enter the holiday season, we wanted to share a quick update on construction progress—and some staff-picked recommendations perfect for group gatherings and gifting for the gamer and non-gamer alike.


Founding Membership 50% Full!

Thanks again for making Tabletop Library exist. It looks like it’s going to be a really great and one-of-a-kind space. The world really needs more places like this as a refuge from the digital world. - Joel

It’s been awesome hearing from new members who are building Tabletop Library with us. We can’t do it without your support, and invite you to be among the first to join the community. Join by December 31, 2025 to lock in the $199 Founding Member price.

Become a member →

In the meantime, we’ll be popping up at the Pickleball Athletic Club on 40th and Telegraph in Oakland to host a board game night on January 7, 2026.

RSVP here


Construction Update: We’re Taking Shape!

Our construction permit was approved last month, and things are moving. Electrical is being installed, and our custom, reclaimed redwood gaming tables (yes—the perfectly sized ones!) are built. We’re still pacing toward a spring 2026 opening and can’t wait to welcome you.


Give the Gift of Membership

A Tabletop Library membership is the perfect gift for the boardgame lover in your life—whether they’re deep in the hobby or just discovering it. Your gift activates when we open in early 2026 and includes:

  • Unlimited access to hundreds of games
  • Exclusive Founding Member perks
  • Early launch-party entry
  • A welcoming community of players ready to meet new friends

Give a gift membership →


Holiday Game Recommendations

Here are games from the Tabletop Library collection that would make great gifts this holiday season (or something to add to your “must play” list for the new year):


🎄 For Families: Flip 7 (Grinch)

A fast, festive push-your-luck game where you flip cards, chase combos, and try not to bust. Simple, silly, and perfect for all ages.

Why we recommend it: Flip 7 is the kind of game you can explain in under five minutes and start playing by the sixth. The Grinch version of the game plays the sameapproachable and full of those small “should I risk one more flip?” moments that kids love and adults can’t resist. 

Buy it now


🃏 For Get-Togethers: The Gang 

A cooperative poker-heist game that’s quick to learn, clever to play, and tense in the best ways.

Why we recommend it: The Gang pulls people in instantly with cooperative decisions, quick rounds, and just enough suspense to keep everyone on the edge of their seat. It’s one of the best “teach-in-3 minutes” games that will warrant a second play right after the first!

Buy it now


☕ For the Foodie: Coffee Rush or Fromage

Coincidentally, both of these games have a racing element, with great table presence and replayability. In Coffee Rush, players collect ingredients and race to fulfill café orders with cute game pieces, such as realistic looking coffee beans, ice cubes, and mint tea leaves—think cozy Overcooked in board-game form. In Fromage, players engage in four mini-games spanning a lazy Susan-style board, crafting award-winning cheese in the French countryside—simple moves, big payoff.

Why we recommend them: These games shine when teaching newcomers: clear goals, tactile components, and playful pressure. Coffee Rush scratches that café-management itch, while Fromage layers familiar mechanics into delightful, rapid-fire mini-rounds. Both are easy to get on the table and instantly charming.

Buy Coffee Rush or Fromage


🖼️ For the Azul Lover: Art Society or The Great Evening Banquet

In Art Society, players bid, curate, and arrange art to build the most elegant gallery—Azul energy with auction flair, and in The Great Evening Banquet players act as event planners who have to seat guests according to their preferences.

Why we recommend them: Art Society blends satisfying spatial arrangement with just enough auction excitement to keep everyone engaged in other players’ turns. The Great Evening Banquet, from beloved Japanese studio Saashi & Saashi, is a standout choice for people who appreciate charming aesthetics and familiar drafting-and-placing tile gameplay.

Buy Art Society or The Great Evening Banquet


💕 For Date Night: High Tide or Tag Team

Fast, head-to-head two-player games that create friendly rivalry and “one more game?” energy.

Why we recommend them: Whether you want lighthearted and highly portable (High Tide) or clever and competitive (Tag Team), these make for easy date-night games. They set up in minutes and deliver that sweet two-player tension without overstaying their welcome.

Buy High Tide or Tag Team


🗂️ For the Collector: Eternal Decks or Jisogi

Beautiful strategic games from Japan

Why we recommend them: These titles are special—not widely available, gorgeously produced, and filled with elegant mechanics. In Eternal Decks, up to four players work together to fulfill patterns on the “board” (a silk handkerchief) before their deck of cards run out. And in Jisogi, players take on the roles of “dead inside” anime studio employees who gather resources to produce the best shows. They’re perfect for the collector who seems to have everything.

Buy Eternal Decks or Jisogi


🗞️ For the Person Who Has It All: Senet subscription

A collectible board game magazine featuring essays, interviews, and lush graphic design—perfect for the gamer who loves the hobby beyond the table.

Why we recommend it: Maybe the person you’re shopping for has Eternal Decks or Jisogi already? Instead of giving someone who has everything another box on the shelf, give a smart magazine with great content they’ll go back to all year long.

Buy it now


Vera Devera

Construction is underway! Become a Founding Member

Opening soon

Our construction permit was approved last week! Building is underway, and we're hoping to open in February 2026. Follow along on Instagram and Facebook.


Become a Founding Member

Now that we have a sense of when we’ll open, we’ve begun early membership enrollment. For a limited time, Founding Memberships are available for $199 per quarter. This grants you:

✓ Unlimited access to our space & 600+ game collection

✓ Reserve tables & join events for free

✓ 5 guest passes per month

✓ 10% discount on food & games

✓ Concierge service to find games and players

As a Founding Member, you will lock in the price for up to 2 years, have your name featured on the Founders plaque, and get an exclusive t-shirt and board game bag in addition to the new members welcome box. Plus you’ll get early access via our launch party!


Curious about our programs?

Wondering what games we'll have? Now you can browse our collection of 400 (and growing) games. They're organized by our take on a dewey decimal system for board games that we call TLCS—if you're curious, here's Nabeel talking about where it came from.

Our mission is to help you feel at home away from home, and master–or discover a new–favorite. Expect programming like:

🀄 Learn-to-Play nights: Quick 10-minute intros and deep-dive teach sessions, from bridge and mahjong to strategic board games

 🚀 Themed game nights, such as “🎩High Society Games” inspired by Downton Abbey and Bridgerton, or “All Aboard: Train Games” that cover the range from family friendly games like Ticket to Ride to Irish Gauge and expert-level games like Carnegie and Shikoku 1889

🏆 Tournaments for TCGs like Magic the GatheringRiftbound, and Pokemon

📖 Organized play for social deduction games like Blood on the Clocktower and roleplaying games like Dungeons & Dragons, Mothership, and Yazeba’s Bed & Breakfast

🎨 Meet-the-creator conversations with the game designers, artists, and authors behind the games we love

🎒 After-school programs and summer camps to teach kids critical thinking skills and good sportsmanship


More about me

Thanks for making it this far! So you’re probably wondering, who’s this Vera Devera? As General Manager, I’m excited to help build the Tabletop Library community from the ground up and it's been a throughline of my career—from helping tech companies launch their first professional networks to growing the Oaklandish Board Gamers from a text thread of five friends into a 500+ member community.

My own boardgaming journey began back in 2014. Until then, my world was mostly ScrabbleCatan, and Monopoly. But once a friend taught me Dominion (and the magic of the “A-B-C: Action, Buy, Cleanup” turn), I was hooked. The clever card interactions, the strategic chaining of actions—it all delivered this perfect dopamine hit that made me want to explore the hobby deeper. 

From there, I moved to Detroit and its vibrant gaming scene deepened my love for everything from cozy card games to heavy strategy titles.

When I returned to the Bay Area in 2018, I created the Petaluma and Oakland Board Gamers groups to help people find gaming friends–after all, games go unplayed if you don’t have anyone to play them with! We grew quickly—too quickly for living-room tables—and the places we did find weren’t comfortable. They were too loud, too dark to read cards, too cold or windy, and we would get glancing looks from the staff and other patrons that telegraphed “you’re taking up too much space and time to play your game.” 

It became clear how much this community needed a space truly designed for play. Joining Tabletop Library feels like coming full circle. The Tabletop Library is a space designed for you and me!

I’m genuinely astounded by the care and intention that Nabeel, Andrew, and David have poured into this space—from the custom-sized tables to the impossibly comfortable chairs. Having spent many game nights in bars with too small tables, or unforgiving brewery stools, I can promise you: this is not that. Tabletop Library is beautiful, inviting, and designed for hours of joyful play.

I’m so excited to be part of this new adventure—and I can’t wait to meet you.

Cheers,

Vera


Andrew Mason

Opening update, and a bit about what TTL will be like

Welcome! We’re still a ways from opening, but I figured some of you might be wondering what's taking so long, so thought we'd check in.

When does Tabletop Library open?

We don’t know! We’re still waiting on some permits. We’re optimistic we’ll have them soon, and the construction team is standing by, so things will move swiftly after that. At this point, we’re guessing this winter. Promise that the next newsletter will include an opening date… or an opening date-ish.

Curious what TTL will be like?

Nabeel and I made a video trying to answer that question. You’ll have to use a little imagination… but hopefully it gives you a sense of what we’re going for.

Other than that, here are a few decisions we’ve mostly made:

  • Hours will be 5pm - 11pm weekdays, 9am - 11pm weekends
  • It’s a membership club - i.e. you’ll pay a quarterly membership fee vs. coming in and paying for a table. We’re finalizing pricing, but the good news is that it’s going to be inline with the expectations of the majority of you who filled out our membership survey.
  • We’ll also have a “coworking” membership tier that will make the space available during the day on weekdays.
  • We’re planning on having events pretty much every night. We’ve concocted an elaborate matrix of board gaming personas and hope to have programming for everyone - for those of you who love discovering new games, for families, for people who like using board games to meet new people, for competitive gamers, etc.
  • Of course you can bring your friends and family to play at TTL, but as many board gamers are painfully aware, some friends and family aren’t super excited to play the latest and greatest historically accurate simulation of peat bog management in mid 17th century East Anglia. So we built an AI concierge text messaging service that helps organize pickup games with other members. You just text, “Hey, I want to play ‘Cabbage Rotation: A Century of Brassica’ this weekend, can you organize a game?” And the concierge will text other members who are interested in this kind of game until it has enough to reserve a table. (Bonus: If your nerdiness transcends cardboard, we did a podcast talking about this and other ways we used AI to help create TTL)

That’s all for now!