Writing

Museum Reciprocal Finder: from schematic to implementation

· side project, JavaScript, PWA, Python

The back of a museum membership card often carries a row of small logos. Each one is a reciprocal network: ASTC for science centers, NARM and ROAM for art and history museums, AZA for zoos and aquariums, AHS for gardens, ACM for children’s museums, and a few more. Show the card at another member institution and you get in free, or at half price, or with a shop discount, depending on the network and sometimes on the museum.

What nobody tells you is where the card actually works. Each network publishes its own list, in its own format, with its own fine print. ASTC won’t let you in free anywhere within 90 miles of your home museum or of where you live. Some NARM museums don’t honor members of institutions close to them. AZA zoos set their discount one zoo at a time. Answering “what does my card get me in Chicago next month?” meant reading several PDFs side by side and doing distance math in my head.

So I built Museum Reciprocal Finder. You tell it which membership you have (or that you have none yet) and where you live, and it sorts museums into Free, ~50% off, Verify and Not covered, with the reason for each. If you don’t have a membership, it finds the cheapest one that covers the kinds of museums you care about. This is the walk from the design doc to the code.

The problem, stated precisely

The app takes three inputs:

  • Membership. One or more home museums, or a set of networks picked directly, or nothing.
  • Home location. A ZIP code or a city. It’s required, because rules like ASTC’s 90 miles are measured from where you live.
  • Search location and radius. Optional, and it defaults to home. This is the “we’re going to Chicago” case: distance rules stay anchored at home while the search moves.

For every museum it returns a tier and a reason: free, discounted, “verify” (the network leaves the benefit to each museum), not covered because the museum isn’t in your networks, or not covered because it’s too close to home.

Two facts from reading the networks’ rules shaped the whole data model:

  1. Benefits and distance rules are set per museum, per network. ACM is a uniform 50% off; AZA varies by zoo; a NARM museum may or may not apply a proximity rule. So each museum stores its own entry for every network it belongs to, and the network supplies defaults only when the museum’s entry is silent.
  2. Distance rules have different anchors. Some measure from your home museum, some from your residence, some between institutions. ASTC uses two anchors with OR logic: being within 90 miles of either your home museum or your home blocks the benefit. That’s why picking a specific home museum gives better answers than picking a network: the app then knows where your home museum is.

The design doc opens with a scope line that saved me from a lot of rabbit holes: this is built primarily for one user, so favor the common path and curated data, and when a rule is fuzzy, a “verify locally” label is good enough.

The schematic

Architecture of Museum Reciprocal Finder

A monthly GitHub Action turns the networks’ listings into committed JSON; a deploy packs and publishes it; the browser caches it and does all the eligibility work locally.

The system has two halves that meet only at a folder of JSON files in the repo.

  • Build time runs once a month in GitHub Actions: collect each network’s list, normalize it to one schema, merge listings of the same museum, geocode, mark well-known museums, validate, and commit data/*.json.
  • Run time is the user’s browser: load the JSON once, cache it with a service worker, and run the eligibility rules and distance math client-side.

There is no server, no database and no API key at runtime, and nothing to pay for between refreshes. The location you type is resolved against lookup tables that ship with the app, so there is no geocoding call when you search either.

Ingestion: one adapter per source

The app covers 14 networks with 13 adapter modules (one web page lists two of the networks). The sources are a mixed bag: participant-list PDFs, a JSON feed behind a museum locator, map data embedded in a page, HTML tables, Google Sites pages and brochure PDFs.

Every adapter splits fetching from parsing. fetch() gets the live document; parse_text(), parse_html() or parse_json() turns a string into records. Because the parsers are pure functions of a string, the tests can feed them saved fixtures, and a manually supplied file can go through exactly the same code. All adapters emit one shape:

@dataclass
class Record:                          # trimmed
    program: str                      # e.g. "ASTC"
    name: str
    city: Optional[str] = None
    state: Optional[str] = None       # 2-letter
    benefit: Optional[str] = None     # None => use the network default
    admits: Optional[int] = None
    tier: Optional[str] = None        # e.g. AZA's "100_or_50"
    exclusion: Optional[dict] = None  # {"radius_mi": 90, "anchors": ["residence"]}
    lat: Optional[float] = None       # when the source already has coordinates
    lng: Optional[float] = None
    raw: str = ""                     # the original line, for auditing

    @property
    def id(self) -> str:
        return slugify(self.name, self.city, self.state)

When a source marks a per-museum rule, the adapter keeps it on the record: ROAM’s 25-mile marker, NARM’s 15 or 50 mile exclusions, the AHS “Local Visitor Exception”, ANCA’s 50-mile note, and AZA’s “100% OR 50%” tier. A few sources list only museum names grouped by state; for those the adapter pins each city by hand.

Sources that opt out of automated access

Two networks, AZA and Time Travelers, say no to automated access in their robots.txt or with a bot challenge. Their adapters carry a BLOCKED_REASON, and the build skips them unless someone has put a file in pipeline/manual_drops/<program>/:

drop = _manual_drop(name)
if not drop and not from_fixtures and getattr(module, "BLOCKED_REASON", None):
    print(f"  {name}: skipped, not scraped ({module.BLOCKED_REASON}); "
          f"add a file to pipeline/manual_drops/{name}/ to include it")
    continue

For those two, I save the current list by hand and drop it in the folder, and the next build writes a plain-text extract beside it. Only the .txt is committed; the original is gitignored, so the repo never republishes the association’s document. A manual drop overrides scraping for any network, which is also my escape hatch if a site changes its terms. AZA’s list runs May to April, so that drop gets replaced each May.

Normalize and merge

The same physical museum shows up in several networks, usually under slightly different names. One record per museum matters, because “the best benefit across all my memberships” only works if a museum’s networks live on one record.

The merge keys on a slug of name, city and state. When a record from another network lands in a city that already has museums, same_museum() decides whether it’s one of them. It compares the distinctive words of the two names (dropping words like “the”, “museum”, “center” and the city’s own name), and it also accepts one name being a word-for-word prefix of the other, as long as the shorter one still has a word of its own:

same_museum("Harbor Science Museum", "Harbor Science Center", "Bayport")      # True
same_museum("Harbor Art Museum", "Harbor Science Museum", "Bayport")          # False

(Fictional museums, real behavior.) Rules only go so far. Museums get renamed, networks keep old names, and one zoo appeared under three different names in three networks. Those cases live in a hand-maintained SAME_MUSEUM alias table, which has about fifty entries now, grouped under comments that say where each came from. Aliases are applied last, so name matching merges everything it can first.

The merge also tags each museum with kinds (art, children’s, science, history, nature, zoo, garden) from regular expressions over its names and its networks’ categories, and repairs or drops broken website links.

Geocoding, offline first

The distance rules are coarse (15 to 90 miles), so a city centroid is accurate enough for nearly every museum. That made geocoding mostly an offline table lookup:

for m in museums:
    if m.get("lat") is not None:          # the source already gave coordinates
        continue
    pt = zips.get(m.get("zip") or "") or (
        places.get(place_key(m["city"], m["state"])) if m.get("city") and m.get("state") else None)
    pt = pt or GEO_OVERRIDES.get(_city_key(m) or "")
    if pt:
        m["lat"], m["lng"] = pt
        continue
    # ...else the cache, else one bounded Nominatim lookup

The order is: coordinates from the source, the ZIP centroid, the Census place and town table, a few hand overrides for places no gazetteer knows (a campus, a valley name), and only then a structured Nominatim query. That last step asks for settlements only (a free-text “Harvard, MA” returns the university), waits 1.1 seconds between requests, caps lookups per run, and writes to a committed cache, so each city is looked up once, ever.

The same Census gazetteers produce the two tables the browser uses: zip_centroids.json (about 34,000 ZIP Code Tabulation Areas) and places.json (every US place and town). They change rarely, so that script runs by hand.

Marking well-known museums

A ★ helps when a search returns two hundred museums. A museum gets one from either of two signals:

  • Wikipedia coverage. A committed Wikidata snapshot counts how many Wikipedia languages cover each US museum, zoo and garden. A museum is starred when its closest-named item within 30 km has 12 or more, so a small museum doesn’t inherit a famous neighbor’s star.
  • Google review volume. A monthly job queries the Places API strictly inside the free tier: at most 900 requests per Pacific-time month (1,000 are free), each recorded before it’s sent so a rerun can’t spend the month twice. It keeps only the place id and the review count rounded down to a bucket. The bar is 10,000 reviews (5,000 for children’s museums, 25,000 for zoos and aquariums).

Validation, the guardrail

A broken parser rarely crashes. Usually it quietly returns fewer rows. So before writing anything, the build compares per-network counts with the previous meta.json:

for prog, prev_n in prev.items():
    now_n = counts.get(prog, 0)
    if prev_n >= 4 and now_n < prev_n * (1 - drop_tolerance):   # 25% by default
        problems.append(f"{prog}: count dropped {prev_n} -> {now_n}; "
                        "likely a broken source; refusing to overwrite good data")

Any problem makes the build exit non-zero. The workflow fails before its commit step, the deploy workflow sees the failure and skips, and last month’s data stays live.

What ships to the browser

The build writes museums.json, and three more files sit next to it: programs.json, zip_centroids.json and the membership catalog. A museum record looks like this (a made-up museum, real shape):

{
  "id": "harbor-science-center-bayport-ca",
  "name": "Harbor Science Center",
  "city": "Bayport", "state": "CA", "zip": "90000",
  "lat": 34.05, "lng": -118.25,
  "kinds": ["science", "children"],
  "programs": {
    "ASTC": { "verified": "2026-10" },
    "ACM":  { "verified": "2026-10", "benefit": "discount_50" },
    "NARM": { "verified": "2026-10",
              "exclusion": { "radius_mi": 15, "anchors": ["home_institution"] } }
  },
  "sources": ["astc", "acm", "narm"],
  "star": true
}

programs.json holds each network’s name, color, defaults and a one-line description for the key. The ASTC default exclusion is the interesting one:

"default_exclusion": {
  "radius_mi": 90,
  "anchors": ["home_institution", "residence"],
  "measure": "straight_line",
  "discretionary": false
}

discretionary: false means the rule always applies. NARM, ROAM and AHS defaults are discretionary: true: the rule applies only where the museum’s own entry states it, which is what the sources mark.

The eligibility engine

The engine was written first, as a pure TypeScript module with no I/O and no framework: data in, classification out. The app itself carries a plain JavaScript port of it in index.html. The heart is the exclusion check, OR across anchors:

function exclusionReason(museum, user, program, { excl, optedIn }) {
  if (!excl) return null;
  if (excl.discretionary && !optedIn) return null;   // optional rule, museum didn't state it
  const rad = excl.radius_mi;
  if (rad == null) return null;
  const homes = user.homeInstitutions.filter(h => h.programs.includes(program));
  for (const anchor of excl.anchors) {
    if (anchor === "home_institution" || anchor === "inter_institution")
      for (const h of homes)
        if (haversineMiles(museum, h) < rad) return { anchor: "home", radius: rad, homeId: h.id };
    if (anchor === "residence" && user.zipCentroid &&
        haversineMiles(museum, user.zipCentroid) < rad) return { anchor: "residence", radius: rad };
  }
  return null;
}

For each museum, the engine walks the networks you hold that the museum belongs to, fills in the benefit and party size from network defaults where needed, and drops the ones a distance rule blocks. What’s left are your options. When there are several, it recommends one: best tier first, then the one that admits the most people, then a stable order.

const TIER_RANK = { free: 3, discount: 2, verify: 1, none: 0 };
const recommended = [...options].sort((a, b) =>
  (TIER_RANK[tierOf(b.benefit)] - TIER_RANK[tierOf(a.benefit)]) ||
  (b.admits - a.admits) || a.program.localeCompare(b.program))[0];

So a children’s museum in both NARM (free for two) and ACM (50% off for up to six) recommends NARM, but the ACM chip stays visible because a family of six might prefer it.

AZA needed one special case. Its reciprocity is in kind: a zoo listed as “100% OR 50%” gives free entry to members of another “100% OR 50%” zoo and half price to everyone else. The engine upgrades the benefit to free when your home zoo (or the AZA tier you ticked) is in that tier.

The classification also returns blocked (your networks that don’t apply here, and why) and others (this museum’s networks you don’t hold), so a card can say “ASTC: within 90 mi of where you live” instead of a bare “not covered”. There is no spatial index: every search runs haversine against all of the roughly 3,200 museums.

Resolving a location

A 5-digit ZIP is looked up in the ZCTA table, which loads on first use. ZIPs with no ZCTA of their own (PO boxes, single buildings) fall back to the mean centroid of their 3-digit prefix, with a note that distances are approximate. A “City, ST” is matched against cities with listed museums, then against the places table, fetched only once someone types a city. A bare city name works if it’s unique; “Springfield” gets a polite request for the state. The same table drives an autocomplete that ranks cities with museums first.

The UI

The whole front end is one index.html with inline CSS and JavaScript, no framework and no build step. The landing page offers three ways in: I’m a member of a museum (search and tick your home museums), I have these associations (tick networks), and I don’t have a membership yet.

Results are cards: a tier badge, colored chips for the networks that apply (the recommended one highlighted), “too close” chips with the reason in a tooltip, and the museum’s kinds. Museums you can visit come first, ordered ASTC, free AZA zoos, ACM, then the rest, nearest first within each group; a chip row reorders by type, by ★, or by distance alone. Cards render 60 at a time. There’s no map; each card links to directions instead.

Finding a membership

The no-membership path grew into its own small tool, with three views: the cheapest memberships, memberships that give you a digital card at checkout (for when you’re standing outside the museum), and museums near you that sell one. A switch flips everything to family levels. Behind it is a catalog of each museum’s cheapest reciprocal level and the networks it includes, read from the museums’ own membership pages.

For the museum types you pick, each membership shows how many museums it covers and its price per museum covered, in cents. Above the list, two panels show the cheapest combination of memberships that reaches 100% and 90% of those museums. With 14 networks there are at most 2^14 network sets, so this is an exact 0/1 knapsack over bitmasks rather than a heuristic:

// reach: network-set bitmask -> cheapest way found to hold exactly that set
let reach = new Map([[0, { cost: 0, items: [] }]]);
for (const g of pool) {
  const mask = maskOf(g.networks), price = priceOf(g);
  for (const [S, v] of [...reach]) {
    const T = S | mask, c = v.cost + price, cur = reach.get(T);
    if (T !== S && (!cur || c < cur.cost)) reach.set(T, { cost: c, items: [...v.items, g] });
  }
}
// then: the cheapest set whose networks cover at least `share` of the picked museums

Every membership card has a “See where it gets you” button that runs the normal results page as if you’d bought it, anchored at that museum and your home.

Prices change, so a workflow opens a reminder issue every January 15 and July 15. The refresh is a manual research pass against the museums’ pages, and a merge script checks that each level only claims networks the museum is listed in and refuses to write if it would drop more than 25% of the catalog.

Offline support

The service worker is about sixty lines, hand-written. It precaches the app shell, icons, font, the packed data files and the ZIP table on install, and deletes old caches on activate. Fetches split by kind:

const path = new URL(e.request.url).pathname;
if (path.includes("/data/") && !/\/(zip_centroids|places)\.json$/.test(path)) {
  // Monthly data: network first so updates land; cache when offline.
  e.respondWith(fetch(e.request).then(res => {
    if (res.ok) { const copy = res.clone(); caches.open(CACHE).then(c => c.put(e.request, copy)); }
    return res;
  }).catch(() => caches.match(e.request)));
  return;
}
// Everything else (shell, ZIP table, places table): cache first.

Requests to other origins (the analytics scripts) pass straight through, and places.json isn’t precached because most people never type a city. When the Census tables are rebuilt, I bump the CACHE version. With the manifest in place, the app installs to a phone home screen and works offline after the first load.

Deployment

The data workflow runs at 06:00 UTC on the first of each month. It runs the pipeline tests first, so a broken parser fails before any data is written, then builds from all 13 adapters (two of them fed only by manual drops) and commits the result if anything changed.

One GitHub Actions detail matters here: pushes made with the workflow’s own GITHUB_TOKEN don’t trigger other push workflows. So the deploy listens for the data workflow to finish, and skips when it failed:

on:
  push:
    branches: [ main ]
  workflow_run:
    workflows: [ "Update reciprocal data" ]
    types: [ completed ]
jobs:
  build-deploy:
    if: github.event_name != 'workflow_run' || github.event.workflow_run.conclusion == 'success'

The app first lived at an address on this site. It now lives at milesmaximizer.com/museum_problem, published by that site’s deploy, which pulls this repo and packs the data on the way out. The old address serves a redirect plus a replacement service worker that deletes the old caches, unregisters itself and reloads open tabs, so installed copies land on the new address.

Testing

  • Adapters: every source has a saved fixture and parser tests against it, so a format change shows up as a failing test rather than as a quiet drop in museums.
  • Build: tests cover merging (renamed museums in the same city, different museums kept apart, aliases, “Saint” vs “St.”), the guardrail, URL cleanup, kinds, and manual drops (a dropped PDF is extracted to text exactly once and never counted twice). The Python suite is a little over a hundred tests.
  • Engine: node:test cases cover haversine (Kansas City to New York is about 1,100 miles), the ASTC rule from both anchors, discretionary rules with and without the museum opting in, the AZA in-kind tier, and overlap recommendations.

Plan versus what I built

The design doc called for React, Vite, Tailwind, Fuse.js for search, a MapLibre map, Workbox for the service worker, rapidfuzz for merging, and a runtime geocoder for street addresses. What shipped is one HTML file with plain JavaScript, substring search, no map, a hand-written service worker, word rules plus an alias table, and ZIP or city input only. The parts of the plan that survived intact were the ones that came from reading the rules: per-museum entries with network defaults, anchors with OR logic, and a guardrail before publishing.

What I’d change next

  1. One engine, not two. The TypeScript engine has the tests, but the app runs a hand port in index.html, and no workflow runs the engine tests. The port should either be generated from the module or tested directly.
  2. Review merge near-misses. The alias table grows by hand whenever a network renames someone. A per-build report of same-city pairs that almost matched would turn that into a quick review.
  3. Better coordinates near a boundary. A museum placed at its city centroid 88 miles from home might really be 92. The app’s “call ahead” note covers it, but precise coordinates for museums close to a 90-mile line would help.
  4. Admission prices. The design doc planned for them. With prices, the overlap recommendation could compare “free for two” with “50% off for six” in dollars.
  5. Canada. NARM and ROAM list Canadian museums, but the app is US-only for now, so the adapters skip non-US rows.
← All writing