<?xml version="1.0" encoding="UTF-8"?><rss xmlns:dc="http://purl.org/dc/elements/1.1/" xmlns:content="http://purl.org/rss/1.0/modules/content/" xmlns:atom="http://www.w3.org/2005/Atom" version="2.0"><channel><title><![CDATA[Tech Magnetics Blog]]></title><description><![CDATA[Tech Magnetics Blog]]></description><link>https://bettechmagnetics.hashnode.dev</link><image><url>https://cdn.hashnode.com/res/hashnode/image/upload/v1593680282896/kNC7E8IR4.png</url><title>Tech Magnetics Blog</title><link>https://bettechmagnetics.hashnode.dev</link></image><generator>RSS for Node</generator><lastBuildDate>Fri, 09 Oct 2026 07:48:16 GMT</lastBuildDate><atom:link href="https://bettechmagnetics.hashnode.dev/rss.xml" rel="self" type="application/rss+xml"/><language><![CDATA[en]]></language><ttl>60</ttl><item><title><![CDATA[Designing a Sports Data Schema: Fixtures, Markets, Outcomes and Odds in PostgreSQL]]></title><description><![CDATA[The schema you regret in month three
Almost every sports app starts the same way. Someone creates a table:
sql
CREATE TABLE matches (
  id serial PRIMARY KEY,
  home_team text,
  away_team text,
  hom]]></description><link>https://bettechmagnetics.hashnode.dev/designing-a-sports-data-schema-fixtures-markets-outcomes-and-odds-in-postgresql</link><guid isPermaLink="true">https://bettechmagnetics.hashnode.dev/designing-a-sports-data-schema-fixtures-markets-outcomes-and-odds-in-postgresql</guid><category><![CDATA[PostgreSQL]]></category><category><![CDATA[database design]]></category><category><![CDATA[SQL]]></category><category><![CDATA[sports data]]></category><category><![CDATA[backend]]></category><dc:creator><![CDATA[Vijay Choudhary]]></dc:creator><pubDate>Wed, 07 Oct 2026 16:30:21 GMT</pubDate><enclosure url="https://cdn.hashnode.com/uploads/covers/6a47ac8dec0b8a17b4dc915d/9f4b3f4b-a501-49a1-84d5-45ca1ae05f07.png" length="0" type="image/jpeg"/><content:encoded><![CDATA[<p>The schema you regret in month three</p>
<p>Almost every sports app starts the same way. Someone creates a table:</p>
<p>sql
CREATE TABLE matches (
  id serial PRIMARY KEY,
  home_team text,
  away_team text,
  home_odds float,
  draw_odds float,
  away_odds float,
  score text
);</p>
<p>It works beautifully for exactly one sprint. Then reality arrives:</p>
<p>Product wants over/under and handicaps. You add over_odds, under_odds, line. Then ah_home, ah_away.
A second bookmaker shows up. Now every column needs a duplicate.
Someone asks for line movement. You've been overwriting prices.
Tennis arrives. There's no "home team," just two players.
Golf arrives. There are 156 competitors in one event.
A player prop market needs a column for the player.
float prices start showing 1.9100000381.</p>
<p>By month three you have a 60-column table, five nullable-by-convention fields, and a migration nobody wants to run.</p>
<p>This post builds the schema you'd wish you'd started with: one that handles 13 sports, any number of markets and bookmakers, live price changes, and full price history, without a column per bet type.</p>
<p>Requirements: PostgreSQL 15+ (for NULLS NOT DISTINCT). Everything else works on older versions with minor changes. This is a general design pattern and a starting point, not a drop-in production system.</p>
<p>Design principles</p>
<p>Before any SQL, agree on the rules. Every decision below follows from these:</p>
<p>Model the world, not the UI. Fixtures, participants, markets, selections, and prices are separate concepts.
Rows, not columns, for variability. A new market type should be a new row, never a new column.
Sport-agnostic core, sport-specific edges. One set of tables serves all sports. Differences live in data and JSONB, not in forked schemas.
Current state and history are different workloads. Keep them in different tables.
Money is numeric, never float.
Never delete facts. Postponed, suspended, and corrected things stay in the database with a status.
Expect change. Schema evolution should be additive.
The mental model
text
sport
  └── competition
        └── fixture (an event)
              ├── fixture_participant ──► participant (team / player / pair)
              └── market
                    └── selection (an outcome you can bet on)
                          └── odds_current  ◄── bookmaker
                          └── odds_history  (append-only, partitioned)</p>
<p>Read it top to bottom as a sentence: a sport has competitions; a competition has fixtures; a fixture has participants and markets; a market has selections; a selection has a price from each bookmaker, now and over time.</p>
<p>Step 1: The reference layer</p>
<p>Small, boring, stable tables. Get these right and everything else is easy.</p>
<p>sql
CREATE TABLE sport (
  id    smallint GENERATED ALWAYS AS IDENTITY PRIMARY KEY,
  slug  text NOT NULL UNIQUE,          -- 'football', 'tennis', 'horse-racing'
  name  text NOT NULL
);</p>
<p>CREATE TABLE competition (
  id            integer GENERATED ALWAYS AS IDENTITY PRIMARY KEY,
  sport_id      smallint NOT NULL REFERENCES sport(id),
  slug          text NOT NULL,
  name          text NOT NULL,
  country_code  char(2),
  UNIQUE (sport_id, slug)
);</p>
<p>CREATE TABLE bookmaker (
  id    integer GENERATED ALWAYS AS IDENTITY PRIMARY KEY,
  slug  text NOT NULL UNIQUE,
  name  text NOT NULL
);</p>
<p>Use GENERATED ALWAYS AS IDENTITY rather than serial. It's the modern, standards-based approach and avoids sequence-ownership surprises.</p>
<p>Participants: teams, players, and pairs in one table</p>
<p>The fastest way to make tennis, golf, and combat sports painful is to create a team table and a separate player table. Instead, use one participant concept.</p>
<p>sql
CREATE TABLE participant (
  id            bigint GENERATED ALWAYS AS IDENTITY PRIMARY KEY,
  sport_id      smallint NOT NULL REFERENCES sport(id),
  kind          text NOT NULL CHECK (kind IN ('team', 'player', 'pair')),
  name          text NOT NULL,
  short_name    text,
  country_code  char(2)
);</p>
<p>CREATE INDEX idx_participant_name ON participant (sport_id, lower(name));</p>
<p>A football club is a team. A tennis player is a player. A doubles pair is a pair. A golfer is a player. The rest of the schema doesn't care.</p>
<p>Step 2: Fixtures and who plays in them</p>
<p>The word fixture (or event) is deliberately generic. A football match, a tennis match, a golf tournament round, a horse race, and a UFC bout are all "things that happen at a time with participants."</p>
<p>sql
CREATE TABLE fixture (
  id              bigint GENERATED ALWAYS AS IDENTITY PRIMARY KEY,
  sport_id        smallint NOT NULL REFERENCES sport(id),
  competition_id  integer  NOT NULL REFERENCES competition(id),
  season          text,
  round           text,
  start_time      timestamptz NOT NULL,
  status          text NOT NULL DEFAULT 'scheduled'
                  CHECK (status IN ('scheduled','live','finished',
                                    'postponed','cancelled','abandoned')),
  neutral_venue   boolean NOT NULL DEFAULT false,
  attributes      jsonb NOT NULL DEFAULT '{}'::jsonb,   -- sport-specific extras
  created_at      timestamptz NOT NULL DEFAULT now(),
  updated_at      timestamptz NOT NULL DEFAULT now()
);</p>
<p>CREATE INDEX idx_fixture_sport_time ON fixture (sport_id, start_time);
CREATE INDEX idx_fixture_comp_time  ON fixture (competition_id, start_time);
CREATE INDEX idx_fixture_live       ON fixture (sport_id) WHERE status = 'live';</p>
<p>That last index is a partial index: it only covers live fixtures, so it stays tiny and fast no matter how many historical rows you accumulate.</p>
<p>Always store start_time as timestamptz. Naive timestamps are the root of countless "the match shows on the wrong day" bugs.</p>
<p>Participants in a fixture: roles, not columns
sql
CREATE TABLE fixture_participant (
  fixture_id      bigint NOT NULL REFERENCES fixture(id) ON DELETE CASCADE,
  participant_id  bigint NOT NULL REFERENCES participant(id),
  role            text   NOT NULL CHECK (role IN ('home', 'away', 'competitor')),
  PRIMARY KEY (fixture_id, participant_id)
);</p>
<p>-- At most one home and one away per fixture
CREATE UNIQUE INDEX uq_fixture_home ON fixture_participant (fixture_id) WHERE role = 'home';
CREATE UNIQUE INDEX uq_fixture_away ON fixture_participant (fixture_id) WHERE role = 'away';</p>
<p>CREATE INDEX idx_fp_participant ON fixture_participant (participant_id);</p>
<p>This one table handles three shapes:</p>
<p>Sport shape	Roles used
Head-to-head with home/away (football, basketball, baseball, hockey)	home + away
Head-to-head, no venue concept (tennis, combat sports, esports)	two competitor rows (or home/away by convention, but be consistent)
Field events (golf, horse racing)	N × competitor</p>
<p>The partial unique indexes enforce "one home, one away" only where it applies. No triggers required.</p>
<p>Step 3: Markets and selections, the heart of the design</p>
<p>This is where naive schemas die. A market is a question you can price. A selection is one possible answer.</p>
<p>Market	Selections
Match winner (1X2)	home, draw, away
Moneyline	home, away
Total goals 2.5	over, under
Handicap −1.5	home, away
Anytime scorer (player X)	yes, no
Tournament winner	one per participant</p>
<p>All of them fit the same two tables.</p>
<p>sql
CREATE TABLE market_type (
  id        smallint GENERATED ALWAYS AS IDENTITY PRIMARY KEY,
  code      text NOT NULL UNIQUE,      -- '1X2','MONEYLINE','TOTAL','SPREAD','PLAYER_PROP','OUTRIGHT'
  has_line  boolean NOT NULL DEFAULT false
);</p>
<p>INSERT INTO market_type (code, has_line) VALUES
  ('1X2', false), ('MONEYLINE', false), ('TOTAL', true),
  ('SPREAD', true), ('PLAYER_PROP', true), ('OUTRIGHT', false);</p>
<p>CREATE TABLE market (
  id                      bigint GENERATED ALWAYS AS IDENTITY PRIMARY KEY,
  fixture_id              bigint   NOT NULL REFERENCES fixture(id) ON DELETE CASCADE,
  market_type_id          smallint NOT NULL REFERENCES market_type(id),
  period                  text     NOT NULL DEFAULT 'full_time',  -- 'full_time','1h','q1','set_1'
  line                    numeric(6,2),                           -- 2.5, -1.5, 22.5 ...
  subject_participant_id  bigint REFERENCES participant(id),      -- player props
  status                  text NOT NULL DEFAULT 'open'
                          CHECK (status IN ('open','suspended','closed','settled')),
  UNIQUE NULLS NOT DISTINCT
    (fixture_id, market_type_id, period, line, subject_participant_id)
);</p>
<p>CREATE INDEX idx_market_fixture ON market (fixture_id);</p>
<p>CREATE TABLE selection (
  id              bigint GENERATED ALWAYS AS IDENTITY PRIMARY KEY,
  market_id       bigint NOT NULL REFERENCES market(id) ON DELETE CASCADE,
  code            text   NOT NULL,       -- 'home','draw','away','over','under','yes','no', or participant id
  participant_id  bigint REFERENCES participant(id),
  UNIQUE (market_id, code)
);</p>
<p>Three subtle choices here deserve explanation:</p>
<p>UNIQUE NULLS NOT DISTINCT. By default, PostgreSQL treats NULLs as distinct, so two "1X2, full_time, NULL line" rows would both be allowed. This PostgreSQL 15 feature makes NULL = NULL for uniqueness, which is exactly what you want for markets without a line.</p>
<p>period as text. Football has halves, basketball quarters, tennis sets. A text code ('1h', 'q3', 'set_2') avoids an enum you'll need to alter every time a sport adds a variation.</p>
<p>status on the market. This is where the four states from market suspension handling live: open, suspended, closed, settled. Boolean is_active columns can't represent that.</p>
<p>Step 4: Prices, current state vs. history</p>
<p>Here's the insight that saves most odds systems: "what's the price right now?" and "what was the price at 20:12:41?" are different workloads.</p>
<p>Current state is read constantly, updated constantly, and small (one row per selection per bookmaker).
History is written constantly, read rarely, and grows without bound.</p>
<p>Mixing them in one table makes both slow. Split them.</p>
<p>sql
CREATE TABLE odds_current (
  selection_id  bigint  NOT NULL REFERENCES selection(id) ON DELETE CASCADE,
  bookmaker_id  integer NOT NULL REFERENCES bookmaker(id),
  price         numeric(9,3) NOT NULL CHECK (price &gt; 1),   -- decimal odds
  state         text NOT NULL DEFAULT 'open' CHECK (state IN ('open','suspended')),
  is_live       boolean NOT NULL DEFAULT false,
  updated_at    timestamptz NOT NULL DEFAULT now(),
  PRIMARY KEY (selection_id, bookmaker_id)
) WITH (fillfactor = 80);</p>
<p>Store decimal odds as the canonical format. American (−110) and fractional (11/10) are display concerns. One canonical representation avoids an entire class of conversion bugs.</p>
<p>fillfactor = 80 leaves room on each page so frequent price updates can be HOT updates (heap-only tuples) that don't touch indexes. This table is update-heavy, so it matters. Notice we deliberately don't index price, because indexing a constantly changing column kills HOT updates.</p>
<p>History: append-only and partitioned
sql
CREATE TABLE odds_history (
  selection_id  bigint       NOT NULL,
  bookmaker_id  integer      NOT NULL,
  captured_at   timestamptz  NOT NULL,
  price         numeric(9,3) NOT NULL,
  state         text         NOT NULL,
  PRIMARY KEY (selection_id, bookmaker_id, captured_at)
) PARTITION BY RANGE (captured_at);</p>
<p>CREATE TABLE odds_history_2026_10 PARTITION OF odds_history
  FOR VALUES FROM ('2026-10-01') TO ('2026-11-01');</p>
<p>CREATE TABLE odds_history_2026_11 PARTITION OF odds_history
  FOR VALUES FROM ('2026-11-01') TO ('2026-12-01');</p>
<p>CREATE INDEX idx_odds_history_time ON odds_history USING brin (captured_at);</p>
<p>Why partition by month?</p>
<p>Retention is a DROP, not a DELETE. Dropping a partition is instant. Deleting 400 million rows is a weekend.
Queries on recent data skip old partitions automatically (partition pruning).
Vacuum and index maintenance stay manageable.
A BRIN index on a time column in an append-only table is tiny and very effective.</p>
<p>Notice there are no foreign keys on odds_history. For a high-write append table, that's a deliberate trade-off: you give up referential checks to avoid per-insert lookups. The write path is the only thing that inserts here, so it's enforceable in code.</p>
<p>Create next month's partition ahead of time with a scheduled job (or the pg_partman extension). A missing partition means failed inserts at midnight on the 1st, a classic.</p>
<p>Step 5: The upsert that writes both tables atomically</p>
<p>Every price update should: update the current row, and append to history, but only if the price or state actually changed. Providers often resend identical values, and storing those as history is pure waste.</p>
<p>PostgreSQL can do this in a single statement:</p>
<p>sql
WITH upsert AS (
  INSERT INTO odds_current (selection_id, bookmaker_id, price, state, is_live, updated_at)
  VALUES (%(selection_id)s, %(bookmaker_id)s, %(price)s, %(state)s, %(is_live)s, %(ts)s)
  ON CONFLICT (selection_id, bookmaker_id) DO UPDATE
     SET price      = EXCLUDED.price,
         state      = EXCLUDED.state,
         is_live    = EXCLUDED.is_live,
         updated_at = EXCLUDED.updated_at
   WHERE (odds_current.price, odds_current.state)
         IS DISTINCT FROM (EXCLUDED.price, EXCLUDED.state)
  RETURNING selection_id, bookmaker_id, price, state, updated_at
)
INSERT INTO odds_history (selection_id, bookmaker_id, captured_at, price, state)
SELECT selection_id, bookmaker_id, updated_at, price, state
FROM upsert
ON CONFLICT DO NOTHING;</p>
<p>The key is the WHERE ... IS DISTINCT FROM on the DO UPDATE. When nothing changed, the update is skipped, RETURNING yields no row, and nothing is appended to history. First-time inserts and real changes flow through to history automatically.</p>
<p>Calling it from Python
python
import psycopg
from datetime import datetime, timezone
from decimal import Decimal</p>
<p>UPSERT_SQL = """ ...the statement above... """</p>
<p>def write_prices(conn: psycopg.Connection, updates: list[dict]) -&gt; None:
    rows = [
        {
            "selection_id": u["selection_id"],
            "bookmaker_id": u["bookmaker_id"],
            "price": Decimal(str(u["price"])),      # never pass floats
            "state": u.get("state", "open"),
            "is_live": u.get("is_live", False),
            "ts": u.get("ts") or datetime.now(timezone.utc),
        }
        for u in updates
        if Decimal(str(u["price"])) &gt; 1
    ]
    with conn.cursor() as cur:
        cur.executemany(UPSERT_SQL, rows)
    conn.commit()</p>
<p>Batch writes. Ten thousand single-row transactions per second will melt a database that handles ten thousand rows in one batch without noticing.</p>
<p>Step 6: Results and corrections</p>
<p>Results are versioned claims, not facts, so the same append-only thinking applies.</p>
<p>sql
CREATE TABLE fixture_result (
  fixture_id   bigint NOT NULL REFERENCES fixture(id) ON DELETE CASCADE,
  version      integer NOT NULL,
  status       text    NOT NULL CHECK (status IN ('provisional','final','abandoned','void')),
  home_score   integer,
  away_score   integer,
  detail       jsonb NOT NULL DEFAULT '{}'::jsonb,   -- sets, innings, finishing order
  received_at  timestamptz NOT NULL DEFAULT now(),
  PRIMARY KEY (fixture_id, version)
);</p>
<p>CREATE VIEW fixture_latest_result AS
SELECT DISTINCT ON (fixture_id) *
FROM fixture_result
ORDER BY fixture_id, version DESC;</p>
<p>The view always exposes the latest version, while the table keeps every correction for audit and re-settlement.</p>
<p>Use detail (JSONB) for the sport-specific shape: tennis set scores, cricket innings, or the finishing order of a race. Keep the common fields as real columns so you can index and query them cheaply.</p>
<p>Step 7: External IDs, the unglamorous essential</p>
<p>You will ingest data from providers who each have their own identifiers. Never put provider IDs on your core tables. Use a crosswalk.</p>
<p>sql
CREATE TABLE external_id (
  entity_type  text   NOT NULL CHECK (entity_type IN
                 ('sport','competition','participant','fixture','market','bookmaker')),
  provider     text   NOT NULL,
  external_id  text   NOT NULL,
  internal_id  bigint NOT NULL,
  PRIMARY KEY (entity_type, provider, external_id)
);</p>
<p>CREATE INDEX idx_external_internal ON external_id (entity_type, internal_id);</p>
<p>This is what lets you swap or add a provider without rewriting your schema, and it's the table that fixture-matching code writes into when it decides two provider records are the same real-world match.</p>
<p>The queries you'll actually run</p>
<p>A schema is only good if the common questions are easy to ask.</p>
<ol>
<li>Best available 1X2 price per upcoming fixture
sql
SELECT f.id,
f.start_time,
max(oc.price) FILTER (WHERE s.code = 'home') AS best_home,
max(oc.price) FILTER (WHERE s.code = 'draw') AS best_draw,
max(oc.price) FILTER (WHERE s.code = 'away') AS best_away</li>
</ol>
<p>FROM fixture f
JOIN market m        ON m.fixture_id = f.id AND m.period = 'full_time'
JOIN market_type mt  ON mt.id = m.market_type_id AND mt.code = '1X2'
JOIN selection s     ON s.market_id = m.id
JOIN odds_current oc ON oc.selection_id = s.id AND oc.state = 'open'
WHERE f.sport_id = (SELECT id FROM sport WHERE slug = 'football')
  AND f.start_time BETWEEN now() AND now() + interval '24 hours'
GROUP BY f.id, f.start_time
ORDER BY f.start_time;</p>
<p>FILTER pivots rows into columns without a pile of CASE expressions, and it only returns open prices, so suspended markets never leak to your UI.</p>
<ol>
<li>Bookmaker margin (overround) for a market</li>
</ol>
<p>The sum of implied probabilities minus 1 tells you how much margin a bookmaker is charging.</p>
<p>sql
SELECT oc.bookmaker_id,
       round(sum(1 / oc.price) - 1, 4) AS overround
FROM odds_current oc
JOIN selection s ON s.id = oc.selection_id
WHERE s.market_id = %(market_id)s
  AND oc.state = 'open'
GROUP BY oc.bookmaker_id
HAVING count(<em>) = (SELECT count(</em>) FROM selection WHERE market_id = %(market_id)s);</p>
<p>The HAVING clause ensures you only compute margin for bookmakers who have prices for every selection. A partial set gives a meaningless number.</p>
<ol>
<li>Line movement for one selection
sql
SELECT captured_at,
price,
price - lag(price) OVER (ORDER BY captured_at) AS change</li>
</ol>
<p>FROM odds_history
WHERE selection_id = %(selection_id)s
  AND bookmaker_id = %(bookmaker_id)s
  AND captured_at &gt; now() - interval '24 hours'
ORDER BY captured_at;</p>
<p>Filtering on captured_at lets PostgreSQL prune to only the relevant partitions.</p>
<ol>
<li>Closing price for every selection in a fixture</li>
</ol>
<p>Closing line value is a standard measure of pricing quality, and it's a single query with DISTINCT ON:</p>
<p>sql
SELECT DISTINCT ON (h.selection_id)
       h.selection_id,
       h.price AS closing_price,
       h.captured_at
FROM odds_history h
JOIN selection s ON s.id = h.selection_id
JOIN market m    ON m.id = s.market_id
JOIN fixture f   ON f.id = m.fixture_id
WHERE m.fixture_id = %(fixture_id)s
  AND h.bookmaker_id = %(bookmaker_id)s
  AND h.captured_at &lt;= f.start_time
ORDER BY h.selection_id, h.captured_at DESC;
5. Implied probability, normalized
sql
WITH raw AS (
  SELECT s.code, 1 / oc.price AS p
  FROM odds_current oc
  JOIN selection s ON s.id = oc.selection_id
  WHERE s.market_id = %(market_id)s
    AND oc.bookmaker_id = %(bookmaker_id)s
)
SELECT code,
       round(p / sum(p) OVER (), 4) AS fair_probability
FROM raw;</p>
<p>Dividing each implied probability by their sum strips out the bookmaker margin to give a no-vig estimate.</p>
<p>Modeling each sport: what actually differs</p>
<p>The core stays identical. Here's what to watch per sport:</p>
<p>Sport	Participants	Modeling notes
Football	Teams, home/away	1X2 (3 selections), draws exist; watch neutral_venue
Basketball	Teams	Moneyline has no draw; use period for quarters, include overtime in totals rules
American football	Teams	Spreads and totals dominate; period for quarters and halves
Cricket	Teams	Multi-day matches: start_time isn't enough, so store innings and format in attributes
Tennis	Players or pairs	No home/away, so use competitor; per-set periods; retirement affects settlement
Baseball	Teams	Doubleheaders: same teams, same day, so start_time must distinguish them; starting pitchers affect markets
Esports	Teams or players	Best-of-N series vs. per-map markets, so use period = 'map_2'
Combat sports	Players	Method-of-victory and round markets; fight cards have many bouts as separate fixtures
Volleyball	Teams	Set-based, so period per set; totals are in points
Handball	Teams	High-scoring totals; halves as periods
Ice hockey	Teams	Regulation vs. overtime/shootout outcomes need explicit period definitions
Golf	Players, N-way field	Outrights and placings; ties and dead-heat rules; one fixture per round or tournament, so decide once
Horse racing	Runners (competitor)	Non-runners, each-way terms, and finishing order in detail; identified by course and race time</p>
<p>Two recurring lessons: put period and line into the market identity, and use JSONB for the genuinely sport-specific payload instead of adding columns.</p>
<p>Performance and operations checklist</p>
<p>A few settings and habits that matter once volume grows:</p>
<p>Tune autovacuum on odds_current. It churns constantly. Lower autovacuum_vacuum_scale_factor for this table specifically:
sql
  ALTER TABLE odds_current SET (
    autovacuum_vacuum_scale_factor = 0.02,
    autovacuum_analyze_scale_factor = 0.02
  );
Keep the hot table narrow. Don't add wide columns to odds_current.
Index for your actual queries, and check with EXPLAIN (ANALYZE, BUFFERS), not intuition.
Pre-create partitions and alert if next month's doesn't exist.
Batch writes (hundreds or thousands of rows per statement).
Use a connection pooler like PgBouncer for streaming workloads.
Cache reads of hot fixtures (Redis or in-process) rather than hammering Postgres for the same scoreboard.
Archive cold partitions to cheaper storage before dropping them if you need history for analytics.
Monitor table bloat, replication lag, and slow queries from day one.
Evolving the schema without breaking everything</p>
<p>The best schemas change by adding, not altering.</p>
<p>Safe (additive)	Risky (breaking)
New table	Renaming a column
New nullable column	Changing a column type
New market_type row	Splitting one table into two
New period code	Removing a status value
New JSONB key	Tightening a CHECK constraint</p>
<p>Treat your own internal views and APIs the same way. This is the same philosophy behind stable versioned APIs: keep the contract steady, add fields freely, and reserve breaking changes for a new version.</p>
<p>Or skip the modeling and consume a normalized feed</p>
<p>Everything above is the work of storing sports data. If you'd rather not design, ingest, and maintain it for 13 sports, you can start from a provider that already ships a normalized schema.</p>
<p>Orbistats covers 13 sports (football, basketball, American football, cricket, tennis, baseball, esports, combat sports, volleyball, handball, ice hockey, golf, and horse racing) behind one consistent API. Its documentation lists resources that line up neatly with the tables you just designed:</p>
<p>Orbistats resource	Maps to in this post
fixtures, results	fixture, fixture_result
competitions, countries	competition, country_code
teams, players	participant
lineups, events	fixture detail / events tables
standings, statistics	derived tables
odds	market, selection, odds_current, odds_history</p>
<p>(The schema in this post is a general design pattern, not a description of Orbistats' internal database.)</p>
<p>Here's which product covers which part:</p>
<p>Sports Data API: fixtures, results, standings, teams, competitions, and countries.
Live Scores API: in-play events like goals, cards, and substitutions.
Sports Statistics API: team and player statistics, including xG and season aggregates.
Odds API: 1X2, moneyline, spreads, handicaps, totals, player props, futures, pre-match and live odds, and line movement, normalized across bookmakers.
Historical Sports Data API: a multi-season archive (stated coverage back to 2010) for backtesting and analytics, which is a natural source for backfilling your partitions.
WebSocket API and Webhooks: two ways to receive live updates without polling.</p>
<p>Building for a particular sport? Start from the coverage pages for Football, Cricket, Tennis, or Horse Racing.</p>
<p>Try it in 10 minutes
Create a free account and grab a key at orbistats.com/signup.
Follow the Quickstart. Auth is a Bearer token against <a href="https://api.orbistats.com/v1/">https://api.orbistats.com/v1/</a>.
Explore responses with no code in the Sandbox.
Check exact field names in the API Reference and Documentation before writing your ingestion mapping.</p>
<p>A minimal fetch to start filling your fixture table:</p>
<p>python
import requests</p>
<p>API_KEY = "YOUR_API_KEY"
BASE = "<a href="https://api.orbistats.com/v1">https://api.orbistats.com/v1</a>"</p>
<p>resp = requests.get(
    f"{BASE}/football/fixtures",
    headers={"Authorization": f"Bearer {API_KEY}"},
    timeout=10,
)
resp.raise_for_status()</p>
<p>for fx in resp.json().get("data", [])[:3]:   # envelope shape: confirm in API Reference
    print(fx)</p>
<p>The free tier is documented at 150 requests per day, plenty to prototype your schema against real data. See pricing when you're ready to scale. New to terms like overround, implied probability, or xG? The Glossary and Guides cover them.</p>
<p>Common mistakes (and the fix)
Mistake	Consequence	Fix
One column per market (over_odds, ah_home...)	Schema change per new bet type	market + selection rows
float prices	Rounding drift, comparison bugs	numeric(9,3) decimal odds
Overwriting prices in place	No line movement, no backtesting	Separate append-only odds_history
Current + history in one table	Slow reads and bloated updates	Split by workload
Separate team and player tables	Tennis, golf, and combat need hacks	Single participant + kind
home_team_id / away_team_id on fixture	Breaks field events and neutral venues	fixture_participant with roles
Boolean is_active on markets	Can't express suspended vs. settled	Explicit status
Naive timestamps	Wrong-day matches	timestamptz everywhere
Provider IDs on core tables	Vendor lock-in	external_id crosswalk
Unpartitioned history	Slow queries, painful deletes	Range partition by time
Writing identical prices repeatedly	Bloated history	Upsert with IS DISTINCT FROM
Hard-deleting postponed fixtures	Broken references, lost audit	Status change, never delete
Production checklist
 Core tables are sport-agnostic, with sport differences in data and JSONB
 One participant table covers teams, players, and pairs
 Fixtures link participants through roles, not fixed columns
 period and line are part of market identity, with NULLS NOT DISTINCT
 Prices are decimal numeric, with other formats derived for display
 odds_current and odds_history are separate tables
 History is partitioned, with partitions created ahead of time
 Writes use a single upsert + history statement, batched
 Results are versioned, with a "latest" view
 Provider IDs live in an external_id crosswalk
 Autovacuum is tuned for the hot table
 Schema changes are additive wherever possible
 Common queries are verified with EXPLAIN (ANALYZE, BUFFERS)
Key takeaways
Model the domain: sports, competitions, fixtures, participants, markets, selections, prices.
Rows, not columns, for variety. A new bet type should never require a migration.
Split current from history. Different workloads deserve different tables.
Decimal odds in numeric are your canonical price format.
Partition history by time and make retention a DROP.
One upsert can update state and append history atomically, skipping no-op updates.
Version results, and expose the latest through a view.
Keep provider IDs out of core tables.
Evolve by adding. Breaking changes are a last resort.
A normalized provider can save you from building and maintaining most of this yourself.
Final thought</p>
<p>A good sports data schema is invisible when it works. New sports, new markets, and new bookmakers become new rows rather than new projects, and the weird edge cases of tennis, golf, and horse racing slot into the same tables as football. Spend a day on the model now, and you'll spend the next year shipping features instead of migrations.</p>
<p>If this helped, share it with the teammate who's about to add over_odds_2_75 as a column. 😄</p>
]]></content:encoded></item><item><title><![CDATA[Settlement, Voids and Result Corrections: Building a Bet Resolution Pipeline That Doesn't Break]]></title><description><![CDATA["The match is over. Just pay the winners."
Every team building a prediction game, fantasy product, or betting-adjacent platform says some version of this sentence in their first planning meeting:
"Onc]]></description><link>https://bettechmagnetics.hashnode.dev/settlement-voids-and-result-corrections-building-a-bet-resolution-pipeline-that-doesn-t-break</link><guid isPermaLink="true">https://bettechmagnetics.hashnode.dev/settlement-voids-and-result-corrections-building-a-bet-resolution-pipeline-that-doesn-t-break</guid><category><![CDATA[Sports API]]></category><category><![CDATA[backend]]></category><category><![CDATA[Python]]></category><category><![CDATA[System Design]]></category><category><![CDATA[PostgreSQL]]></category><dc:creator><![CDATA[Vijay Choudhary]]></dc:creator><pubDate>Wed, 07 Oct 2026 16:26:25 GMT</pubDate><enclosure url="https://cdn.hashnode.com/uploads/covers/6a47ac8dec0b8a17b4dc915d/70bf2d42-a193-4152-9d81-559d707cbf2c.png" length="0" type="image/jpeg"/><content:encoded><![CDATA[<p>"The match is over. Just pay the winners."</p>
<p>Every team building a prediction game, fantasy product, or betting-adjacent platform says some version of this sentence in their first planning meeting:</p>
<p>"Once the match ends, we compare the result with each bet and pay out. How hard can it be?"</p>
<p>Then production teaches you otherwise:</p>
<p>The final whistle blows, but the provider later credits the goal to a different player, flipping every player-prop bet.
A match is abandoned at 78'. Which bets are void, and which still stand?
A tennis player retires mid-match.
Two horses dead-heat for first.
An Over 2.75 goals bet lands on exactly 3.
A result you already settled gets corrected 40 minutes later, and money has already moved.</p>
<p>Settlement is where sports data stops being information and becomes money. Bugs here don't produce ugly UI. They produce wrong balances, support floods, and in regulated contexts, compliance problems.</p>
<p>This guide walks through how to design a settlement pipeline that handles the messy cases correctly, with working code.</p>
<p>Note: Settlement rules differ by operator, product, and jurisdiction. The logic below is a technical pattern, not legal or compliance advice. Always encode your published rules, and consult the relevant regulations if real money is involved.</p>
<p>The mental model: settlement is a state machine over versions of a result</p>
<p>The biggest conceptual shift is this: a result is not a fact, it's a claim that can be revised.</p>
<p>Think of the lifecycle like this:</p>
<p>text
 Match live
     ↓
 Result PROVISIONAL   ← scoreboard says "full time", but not official
     ↓
 Result FINAL         ← safe to settle
     ↓
 Bets SETTLED         ← money moves
     ↓
 Result CORRECTED?    ← new version arrives
     ↓
 RE-SETTLEMENT        ← compute the difference, never overwrite history</p>
<p>Two rules fall out of this immediately:</p>
<p>Don't settle on provisional results. "Full time" on a live feed isn't always official.
Never mutate settled records. Append corrections. Keep the audit trail.
The outcomes you actually need to support</p>
<p>Most teams model won and lost, then bolt on the rest in a panic. Model the full set up front.</p>
<p>Outcome	Meaning	Payout (stake S, decimal odds O)
Won	Selection succeeded	S × O
Lost	Selection failed	0
Void	Bet cancelled, stake returned	S
Push	Exact line hit (e.g., Over 3.0 lands on 3), stake returned	S
Half won	Quarter-line, half the stake wins	S/2 × O + S/2
Half lost	Quarter-line, half the stake loses	S/2
Dead heat	Tie for a winning position, stake split	(S ÷ k) × O</p>
<p>A push is functionally a void in payout terms, but keep them as separate labels. Reporting, user messaging, and disputes all care about the difference.</p>
<p>Step 1: Use Decimal, never float</p>
<p>Money and floats don't mix. 0.1 + 0.2 != 0.3 is funny until it's your ledger.</p>
<p>python
from decimal import Decimal, ROUND_HALF_UP
from enum import Enum</p>
<p>CENT = Decimal("0.01")</p>
<p>class Outcome(str, Enum):
    WON       = "won"
    LOST      = "lost"
    VOID      = "void"
    HALF_WON  = "half_won"
    HALF_LOST = "half_lost"</p>
<p>def payout(stake: Decimal, odds: Decimal, outcome: Outcome) -&gt; Decimal:
    if outcome is Outcome.WON:
        amount = stake * odds
    elif outcome is Outcome.LOST:
        amount = Decimal("0")
    elif outcome is Outcome.VOID:
        amount = stake
    elif outcome is Outcome.HALF_WON:
        amount = (stake / 2) * odds + (stake / 2)
    elif outcome is Outcome.HALF_LOST:
        amount = stake / 2
    else:
        raise ValueError(f"unknown outcome: {outcome}")
    return amount.quantize(CENT, rounding=ROUND_HALF_UP)</p>
<p>Always construct decimals from strings (Decimal("1.91")), never from floats (Decimal(1.91) bakes in binary error).</p>
<p>Step 2: Settle totals, including quarter lines and pushes</p>
<p>An Over/Under on a whole or half line is easy. The interesting cases are exact hits (push) and quarter lines like 2.75, which split your stake across two adjacent lines.</p>
<p>python
def _settle_single_line(side: str, line: Decimal, total: Decimal) -&gt; Outcome:
    if total == line:
        return Outcome.VOID                         # push: stake returned
    went_over = total &gt; line
    return Outcome.WON if (side == "over") == went_over else Outcome.LOST</p>
<p>def settle_over_under(side: str, line: Decimal, total: Decimal) -&gt; Outcome:
    is_quarter = (line * 2) % 1 != 0                # 2.25, 2.75, etc.
    if not is_quarter:
        return _settle_single_line(side, line, total)</p>
<pre><code># Quarter line = two half-stake bets on adjacent lines
lo, hi = line - Decimal("0.25"), line + Decimal("0.25")
a = _settle_single_line(side, lo, total)
b = _settle_single_line(side, hi, total)

pair = {a, b}
if pair == {Outcome.WON}:                      return Outcome.WON
if pair == {Outcome.LOST}:                     return Outcome.LOST
if pair == {Outcome.WON,  Outcome.VOID}:       return Outcome.HALF_WON
if pair == {Outcome.LOST, Outcome.VOID}:       return Outcome.HALF_LOST
raise AssertionError(f"impossible combination: {a}, {b}")
</code></pre>
<p>Sanity check: Over 2.75, total = 3. The 2.5 half wins, the 3.0 half pushes, so the result is HALF_WON. Write a unit test for this exact case, because it's the one customers argue about.</p>
<p>Step 3: A decision function with explicit void rules</p>
<p>Voids are where most settlement bugs live, because rules differ by market and sport. Put every rule in one auditable place instead of scattering if statements across workers.</p>
<p>python
from dataclasses import dataclass</p>
<p>VOID_STATUSES = {"cancelled", "abandoned", "postponed", "walkover"}</p>
<p>@dataclass(frozen=True)
class Result:
    match_id: str
    version: int            # increments on every correction
    status: str             # "provisional" | "final" | "abandoned" | ...
    home: int
    away: int</p>
<p>@dataclass(frozen=True)
class Bet:
    id: str
    match_id: str
    market: str             # "1X2" | "TOTAL"
    selection: str          # "home"/"draw"/"away" or "over"/"under"
    line: Decimal | None
    odds: Decimal
    stake: Decimal</p>
<p>def decide(bet: Bet, r: Result) -&gt; Outcome | None:
    """Return an Outcome, or None if the bet is not yet settleable."""
    if r.status == "provisional":
        return None                                  # never settle on provisional</p>
<pre><code>if r.status in VOID_STATUSES:
    return Outcome.VOID                          # encode YOUR published rules here

if r.status != "final":
    return None

if bet.market == "1X2":
    winner = "home" if r.home &gt; r.away else "away" if r.away &gt; r.home else "draw"
    return Outcome.WON if bet.selection == winner else Outcome.LOST

if bet.market == "TOTAL":
    return settle_over_under(bet.selection, bet.line, Decimal(r.home + r.away))

raise NotImplementedError(f"market {bet.market} not supported")
</code></pre>
<p>Notice the function returns None for "not ready". Unknown or unsupported must never default to lost. Leaving a bet pending is recoverable. Wrongly settling it is not.</p>
<p>Sport-specific void and edge-case rules worth documenting</p>
<p>Your published rules need explicit answers to these. They're the cases that cause disputes:</p>
<p>Sport / situation	Question you must answer
Football, abandoned match	Do bets stand if a result market was already decided? Time-based markets?
Tennis, retirement	Settle on completed sets, or void? What if one set is finished?
Cricket, rain / DLS	Which result applies, and what about minimum overs?
Baseball, rain delay	How many innings make a game official?
Golf, ties	Dead-heat reduction, and each-way place terms
Horse racing, non-runners	Price adjustments, dead heats, stewards' enquiries
Combat sports, bout changes	Fighter replaced? Fight cancelled? Draw or no-contest?
Player props	Player didn't play: void or lose?</p>
<p>Write the decisions down before writing code. Code should implement a rulebook, not invent one.</p>
<p>Step 4: Dead heats</p>
<p>Dead heats appear in horse racing, golf, athletics, and any "finishing position" market. The convention: your stake is divided among the dead-heaters, and the winning portion is paid at full odds.</p>
<p>python
def dead_heat_payout(stake: Decimal, odds: Decimal, tied: int) -&gt; Decimal:
    """Selection tied for a winning position with <code>tied</code> total participants."""
    if tied &lt; 1:
        raise ValueError("tied must be &gt;= 1")
    winning_part = stake / tied
    return (winning_part * odds).quantize(CENT, rounding=ROUND_HALF_UP)</p>
<h1>£10 at 5.00, two-way dead heat → pays 25.00 (stake 5 wins at 5.0; other 5 is lost)</h1>
<p>assert dead_heat_payout(Decimal("10"), Decimal("5.00"), 2) == Decimal("25.00")
Step 5: The ledger, where corrections stop being scary</p>
<p>This is the part that separates a toy from a production system.</p>
<p>The wrong way: update the bet row from won to lost and adjust a user balance.</p>
<p>The right way: an append-only ledger. Every settlement or correction is a new row. The balance is derived from the rows. History is never rewritten.</p>
<p>sql
CREATE TABLE results (
    match_id     TEXT        NOT NULL,
    version      INT         NOT NULL,
    status       TEXT        NOT NULL,     -- provisional | final | abandoned ...
    home_score   INT,
    away_score   INT,
    payload      JSONB,
    received_at  TIMESTAMPTZ NOT NULL DEFAULT now(),
    PRIMARY KEY (match_id, version)
);</p>
<p>CREATE TABLE bets (
    id          TEXT PRIMARY KEY,
    match_id    TEXT          NOT NULL,
    market      TEXT          NOT NULL,
    selection   TEXT          NOT NULL,
    line        NUMERIC(6,2),
    odds        NUMERIC(10,4) NOT NULL,
    stake       NUMERIC(14,2) NOT NULL,
    status      TEXT          NOT NULL DEFAULT 'pending'
);</p>
<p>CREATE TABLE ledger (
    id             BIGSERIAL PRIMARY KEY,
    bet_id         TEXT          NOT NULL REFERENCES bets(id),
    result_version INT           NOT NULL,
    outcome        TEXT          NOT NULL,
    amount         NUMERIC(14,2) NOT NULL,   -- signed delta vs. what was already credited
    reason         TEXT,
    created_at     TIMESTAMPTZ   NOT NULL DEFAULT now(),
    UNIQUE (bet_id, result_version)          -- idempotency guard
);</p>
<p>CREATE INDEX idx_ledger_bet ON ledger (bet_id);</p>
<p>That UNIQUE (bet_id, result_version) constraint is doing a lot of work. If a worker crashes and retries, or a webhook is delivered twice, the second insert simply fails. Double settlement becomes physically impossible.</p>
<p>Step 6: Re-settlement as a delta</p>
<p>When a corrected result arrives, compute what the bet should have paid under the new version, subtract what's already been credited, and record the difference.</p>
<p>python
class Ledger:
    """In-memory sketch; in production this is the table above inside a DB transaction."""
    def <strong>init</strong>(self):
        self.rows: list[dict] = []</p>
<pre><code>def credited(self, bet_id: str) -&gt; Decimal:
    return sum((r["amount"] for r in self.rows if r["bet_id"] == bet_id), Decimal("0"))

def has(self, bet_id: str, version: int) -&gt; bool:
    return any(r["bet_id"] == bet_id and r["version"] == version for r in self.rows)

def append(self, bet_id, version, outcome, amount, reason):
    self.rows.append(dict(bet_id=bet_id, version=version,
                          outcome=outcome, amount=amount, reason=reason))
</code></pre>
<p>def settle_or_resettle(bet: Bet, result: Result, ledger: Ledger):
    if ledger.has(bet.id, result.version):
        return None                                   # idempotent: already handled</p>
<pre><code>outcome = decide(bet, result)
if outcome is None:
    return None                                   # not settleable yet

new_total = payout(bet.stake, bet.odds, outcome)
delta = new_total - ledger.credited(bet.id)

reason = "initial settlement" if not ledger.credited(bet.id) and result.version == 1 \
         else f"correction to result v{result.version}"
ledger.append(bet.id, result.version, outcome.value, delta, reason)
return outcome, delta
</code></pre>
<p>Walk through a real scenario:</p>
<p>python
bet = Bet("b1", "match_50231", "TOTAL", "over", Decimal("2.5"),
          Decimal("1.90"), Decimal("10.00"))
ledger = Ledger()</p>
<h1>v1: provider says 2-1, final → total 3 → Over wins</h1>
<p>settle_or_resettle(bet, Result("match_50231", 1, "final", 2, 1), ledger)
print(ledger.credited("b1"))      # 19.00</p>
<h1>v2: a goal is later disallowed → 1-1, total 2 → Over loses</h1>
<p>settle_or_resettle(bet, Result("match_50231", 2, "final", 1, 1), ledger)
print(ledger.credited("b1"))      # 0.00   (delta of -19.00 recorded, history intact)</p>
<h1>Delivering v2 twice is harmless</h1>
<p>settle_or_resettle(bet, Result("match_50231", 2, "final", 1, 1), ledger)
print(ledger.credited("b1"))      # still 0.00</p>
<p>You now have a full audit trail: initial settlement +19.00, correction −19.00, with timestamps and reasons. Finance, support, and compliance all need exactly this.</p>
<p>Step 7: Wire it to real-time result updates</p>
<p>Corrections can arrive at any time, so polling once after full time isn't enough. You need a continuous result watcher.</p>
<p>python
from fastapi import FastAPI, Request, HTTPException</p>
<p>app = FastAPI()</p>
<p>@app.post("/webhook/results")
async def result_webhook(request: Request):
    p = await request.json()
    try:
        result = Result(
            match_id=p["match_id"],
            version=int(p["version"]),        # illustrative, derive if provider lacks one
            status=p["status"],
            home=int(p["home_score"]),
            away=int(p["away_score"]),
        )
    except (KeyError, ValueError):
        raise HTTPException(status_code=400, detail="bad payload")</p>
<pre><code>save_result(result)                       # insert into `results` (PK blocks replays)
enqueue_settlement(result.match_id)       # settle asynchronously, ack fast
return {"ok": True}
</code></pre>
<p>If your provider doesn't send an explicit version number, derive one. Hash the meaningful fields (status, scores, key events) and bump your own version whenever the hash changes. That turns any feed into a versioned one.</p>
<p>python
import hashlib, json</p>
<p>def fingerprint(p: dict) -&gt; str:
    core = {k: p.get(k) for k in ("status", "home_score", "away_score", "events")}
    return hashlib.sha256(json.dumps(core, sort_keys=True).encode()).hexdigest()</p>
<p>And add a safety-net re-check job that re-pulls results for recently settled matches over a defined window (hours or days, depending on the sport) and compares fingerprints. Webhooks are great, but belt and braces is cheap.</p>
<p>Step 8: Confidence gates before money moves</p>
<p>For higher-stakes products, add a gate between "provider says final" and "we pay":</p>
<p>Settlement delay: wait a short, sport-appropriate buffer after final to let early corrections land.
Cross-check: compare the final result against a second source or against live event totals (goals scored in events should equal the final score).
Anomaly flag: scores that look implausible (a 9-8 football match) go to manual review.
Manual override with audit: operators can force-void or force-settle, with a mandatory reason, written to the ledger like everything else.</p>
<p>Event-level consistency checks catch a surprising amount of garbage:</p>
<p>python
def events_match_score(events: list[dict], home: int, away: int) -&gt; bool:
    goals_home = sum(1 for e in events if e["type"] == "goal" and e["side"] == "home")
    goals_away = sum(1 for e in events if e["type"] == "goal" and e["side"] == "away")
    return (goals_home, goals_away) == (home, away)
Testing a settlement engine properly</p>
<p>Settlement code deserves more tests than almost anything else you'll write.</p>
<p>Table-driven unit tests for every market and edge case:</p>
<p>python
import pytest</p>
<p>@pytest.mark.parametrize("side,line,total,expected", [
    ("over",  "2.5",  3, Outcome.WON),
    ("over",  "2.5",  2, Outcome.LOST),
    ("over",  "3.0",  3, Outcome.VOID),       # push
    ("over",  "2.75", 3, Outcome.HALF_WON),
    ("under", "2.75", 3, Outcome.HALF_LOST),
    ("under", "2.25", 2, Outcome.HALF_WON),
])
def test_totals(side, line, total, expected):
    assert settle_over_under(side, Decimal(line), Decimal(total)) == expected</p>
<p>Property tests: for any bet, settling the same result version twice never changes the balance. For any sequence of corrections, the ledger sum equals the payout under the latest version.</p>
<p>Replay tests: run the engine over a season of historical matches and assert aggregate results against known figures. Historical datasets with closing odds are ideal for this kind of regression suite.</p>
<p>Chaos tests: kill the worker mid-batch, deliver duplicates, deliver versions out of order, and confirm the ledger ends in the same state.</p>
<p>Common mistakes (and the fix)
Mistake	Consequence	Fix
Settling on "full time" from a live feed	Paying out on provisional results	Require an explicit final state
Using float for money	Cent-level drift, reconciliation failures	Decimal end to end
Overwriting bet status on correction	No audit trail, can't explain balances	Append-only ledger with deltas
No idempotency key	Duplicate payouts on retry	UNIQUE (bet_id, result_version)
Defaulting unknown cases to lost	Customers lose money to your gaps	Return None, keep pending, alert
Treating push as win/loss	Wrong stake handling	Model void/push explicitly
Ignoring quarter lines	Half-win/half-loss disputes	Split into two adjacent lines
Settling once, never re-checking	Corrections silently ignored	Webhooks plus a re-check window
Rules buried in code	Can't answer "why was this voided?"	Central rulebook, documented publicly
Where the data comes from, and why a unified feed helps</p>
<p>A settlement pipeline is only as good as the results feeding it. You need final scores, event-level detail to cross-check them, notification when something changes, and historical data to test against.</p>
<p>Orbistats offers these pieces under one normalized schema across 13 sports (football, basketball, American football, cricket, tennis, baseball, esports, combat sports, volleyball, handball, ice hockey, golf, and horse racing). Here is how they map to the pipeline above:</p>
<p>Final results and fixtures: The Sports Data API provides fixtures, results, standings, teams, and competitions. This is the source of truth for Step 3.
Event-level cross-checks: The Live Scores API delivers goals, cards, and substitutions, which feed the consistency check in Step 8.
Change notifications: Webhooks push events to your server, matching the pattern in Step 7. For live streaming, the WebSocket API keeps a persistent connection open.
Odds at placement and close: The Odds API covers 1X2, moneyline, spreads, handicaps, totals, player props, and futures in a normalized format, and line movement helps you analyze how prices evolved.
Test and backtest data: The Historical Sports Data API is described as a multi-season archive (stated coverage back to 2010) with fixtures, results, statistics, lineups, match events, and closing odds, which is exactly what a replay test suite needs.</p>
<p>Sports with unusual settlement semantics have dedicated coverage pages worth reading before you write their rules: Horse Racing, Golf, Tennis, and Combat Sports. If you're building for operators, see the Sportsbooks &amp; Trading solution page.</p>
<p>Try it in 10 minutes
Get a free API key at orbistats.com/signup.
Follow the Quickstart. Auth is a Bearer token against <a href="https://api.orbistats.com/v1/">https://api.orbistats.com/v1/</a>.
Explore responses without writing code in the Sandbox.
Confirm exact field names for results and events in the API Reference and Documentation. The payload fields in this post are illustrative.
python
import requests</p>
<p>API_KEY = "YOUR_API_KEY"
BASE = "<a href="https://api.orbistats.com/v1">https://api.orbistats.com/v1</a>"</p>
<p>resp = requests.get(
    f"{BASE}/football/fixtures",
    headers={"Authorization": f"Bearer {API_KEY}"},
    timeout=10,
)
resp.raise_for_status()
print(resp.json())</p>
<p>The free tier is documented at 150 requests/day, plenty for prototyping a settlement engine against real fixtures. See pricing when you're ready to scale. Unfamiliar terms like implied probability or line movement are explained in the Glossary and Guides.</p>
<p>Production checklist
 Results carry an explicit status and version
 Nothing settles on provisional data
 Money uses Decimal end to end
 Every outcome is modeled: won, lost, void, push, half-won, half-lost, dead heat
 All void rules live in one documented rulebook
 The ledger is append-only with UNIQUE (bet_id, result_version)
 Corrections are applied as deltas, never overwrites
 Unknown cases stay pending and alert, never default to lost
 A re-check window catches corrections webhooks miss
 A manual override exists, with mandatory reason and audit
 Tests cover quarter lines, pushes, duplicates, out-of-order versions, and crashes
Key takeaways
A result is a versioned claim, not a permanent fact.
Only settle on final results, and design for corrections from day one.
Model every outcome type, especially pushes, quarter lines, and dead heats.
Use an append-only ledger with idempotency keys so retries and replays are safe.
Apply corrections as deltas and keep the full audit trail.
Unknown never means lost. Fail pending, not wrong.
Write the rulebook before the code, and test it relentlessly.
Reliable results, events, and history from a single normalized source make cross-checks and replay testing far easier.
Final thought</p>
<p>Anyone can pay out when a favorite wins by three goals. A settlement pipeline earns its keep on the weird afternoon when a goal is chalked off, a race ends in a dead heat, a match is abandoned, and the provider revises the score twice. Build for that afternoon, and every ordinary day takes care of itself.</p>
<p>If this helped, share it with the teammate who's about to write UPDATE bets SET status='lost'. 😄</p>
]]></content:encoded></item><item><title><![CDATA[Market Suspensions Explained: Why Odds Vanish Mid-Match and How to Handle It in Code]]></title><description><![CDATA[The ticket that always arrives at the worst moment
It's the 72nd minute. Your app shows Home 1.91 / Draw 3.40 / Away 4.20. A user taps "Home" to add it to their bet slip.
Meanwhile, 400 milliseconds e]]></description><link>https://bettechmagnetics.hashnode.dev/market-suspensions-explained-why-odds-vanish-mid-match-and-how-to-handle-it-in-code</link><guid isPermaLink="true">https://bettechmagnetics.hashnode.dev/market-suspensions-explained-why-odds-vanish-mid-match-and-how-to-handle-it-in-code</guid><category><![CDATA[Odds API]]></category><category><![CDATA[websockets]]></category><category><![CDATA[sports data]]></category><category><![CDATA[Python]]></category><category><![CDATA[API Design]]></category><dc:creator><![CDATA[Vijay Choudhary]]></dc:creator><pubDate>Wed, 07 Oct 2026 16:21:54 GMT</pubDate><enclosure url="https://cdn.hashnode.com/uploads/covers/6a47ac8dec0b8a17b4dc915d/86fb46ae-90e5-4611-a7c4-393617276fb0.png" length="0" type="image/jpeg"/><content:encoded><![CDATA[<p>The ticket that always arrives at the worst moment</p>
<p>It's the 72nd minute. Your app shows Home 1.91 / Draw 3.40 / Away 4.20. A user taps "Home" to add it to their bet slip.</p>
<p>Meanwhile, 400 milliseconds earlier, the striker was brought down in the box.</p>
<p>By the time the tap lands, the price on screen is a lie. If your UI lets the user act on it, you now have one of three problems:</p>
<p>A rejected bet and an angry user
An accepted bet at a price that no longer exists
A support ticket titled "odds disappeared!!"</p>
<p>Welcome to market suspensions: the single most misunderstood part of working with live odds data.</p>
<p>This guide explains what they are, why they happen, what providers actually send you, and how to handle them correctly in code.</p>
<p>What is a market suspension?</p>
<p>A market is a single thing you can price: Match Winner (1X2), Over/Under 2.5, Next Goal, Player to Score, and so on.</p>
<p>A suspension means: this market is temporarily not tradable. Don't show the price as bettable.</p>
<p>Crucially, suspended is not the same as gone. The market usually comes back, often within seconds, frequently with new odds. Treating a temporary suspension as a permanent deletion is one of the most common bugs in odds integrations.</p>
<p>The four states you must distinguish
State	Meaning	Expected next step
Open	Prices are live and tradable	May suspend at any moment
Suspended	Temporarily frozen, prices untrustworthy	Usually reopens, often with new odds
Closed	No more betting on this market	Awaits settlement
Settled	Outcome decided	Terminal, never reopens</p>
<p>If your database has a single boolean is_available, you've already lost information you'll need later.</p>
<p>Why do odds suspend mid-match?</p>
<p>Bookmakers and trading systems suspend markets whenever the real-world situation changes faster than the price can be re-evaluated. The triggers differ by sport.</p>
<p>Football
A goal was scored (or is being checked)
VAR review in progress
A penalty awarded
A red card
A serious injury or stoppage
Basketball
Timeouts and quarter breaks
Technical fouls or reviews
Injury to a key player
Tennis
Break point or game point approaching
Medical timeout
Challenge review
Cricket
Rain delay or bad light
A wicket falling, a review (DRS) in progress
Innings break
Baseball and ice hockey
Weather delays
Goal or run reviews
Pitching or goalie changes
Across every sport
Latency concerns: the trader isn't confident the feed is current
Data feed interruptions
Risk management: unusual betting patterns on a market
Pre-match to in-play transition at kickoff</p>
<p>The pattern is always the same: something just happened, or is about to happen, and the price might be wrong.</p>
<p>What providers actually send you</p>
<p>Providers signal suspension in different ways, and a robust integration has to handle all of them:</p>
<p>An explicit status field ("open", "suspended", "closed")
A boolean flag on the market or each selection
Odds set to null or zero while suspended
The market disappearing from the next snapshot
Silence: no update for several seconds (a de facto suspension)</p>
<p>Here's an illustrative payload. Check your provider's documentation for the exact schema.</p>
<p>json
{
  "match_id": "match_50231",
  "market": "1X2",
  "status": "suspended",
  "seq": 18422,
  "updated_at": "2026-10-15T20:12:41.380Z",
  "odds": {
    "home": 1.91,
    "draw": 3.40,
    "away": 4.20
  }
}</p>
<p>Note that even while suspended, the last-known prices are present. They must not be displayed as live. They're only useful for history and for showing "price before suspension."</p>
<p>The core rule: never show a price you can't stand behind</p>
<p>Build every odds feature around one function:</p>
<p>"Is this price safe to show as tradable right now?"</p>
<p>A price is tradable only if all three are true:</p>
<p>The market status is open
The last update is fresh (within your staleness threshold)
The connection to the feed is healthy</p>
<p>Miss any one and you'll eventually display an invalid price.</p>
<p>Step 1: Model market state properly</p>
<p>Start with an explicit state machine, not a boolean.</p>
<p>python
import time
from dataclasses import dataclass, field
from enum import Enum
from typing import Optional</p>
<p>class MarketState(str, Enum):
    OPEN = "open"
    SUSPENDED = "suspended"
    CLOSED = "closed"
    SETTLED = "settled"</p>
<h1>Which transitions are legal</h1>
<p>ALLOWED = {
    MarketState.OPEN:      {MarketState.OPEN, MarketState.SUSPENDED, MarketState.CLOSED},
    MarketState.SUSPENDED: {MarketState.OPEN, MarketState.SUSPENDED, MarketState.CLOSED},
    MarketState.CLOSED:    {MarketState.CLOSED, MarketState.SETTLED},
    MarketState.SETTLED:   {MarketState.SETTLED},   # terminal
}</p>
<p>@dataclass
class Market:
    match_id: str
    market: str
    state: MarketState = MarketState.SUSPENDED   # fail safe: unknown = not tradable
    odds: dict = field(default_factory=dict)
    seq: int = -1
    updated_at: float = 0.0                      # local monotonic-ish receive time</p>
<p>Notice the default: unknown markets start as suspended. If you've never heard about a market, you shouldn't show it as bettable. Fail closed, not open.</p>
<p>Step 2: Apply updates safely (ordering and idempotency)</p>
<p>WebSocket messages can arrive duplicated or out of order, especially after reconnects. Use a sequence number where available and ignore anything older than what you've seen.</p>
<p>python
class MarketBook:
    def <strong>init</strong>(self, stale_after: float = 5.0):
        self.markets: dict[tuple[str, str], Market] = {}
        self.stale_after = stale_after</p>
<pre><code>def apply(self, msg: dict) -&gt; bool:
    """Apply one update. Returns True if state changed."""
    key = (msg["match_id"], msg["market"])
    new_state = MarketState(msg["status"])
    seq = msg.get("seq", -1)

    cur = self.markets.get(key)
    if cur is None:
        cur = self.markets[key] = Market(*key)

    # 1. Drop duplicates and out-of-order messages
    if seq != -1 and seq &lt;= cur.seq:
        return False

    # 2. Enforce legal transitions (never reopen a settled market)
    if new_state not in ALLOWED[cur.state]:
        return False

    changed = (new_state != cur.state) or (msg.get("odds") != cur.odds)

    cur.state = new_state
    cur.seq = seq if seq != -1 else cur.seq
    cur.updated_at = time.monotonic()
    if msg.get("odds"):
        cur.odds = msg["odds"]
    return changed

def tradable_odds(self, match_id: str, market: str) -&gt; Optional[dict]:
    """The ONLY function the UI/bet-slip should call."""
    m = self.markets.get((match_id, market))
    if m is None or m.state is not MarketState.OPEN:
        return None
    if time.monotonic() - m.updated_at &gt; self.stale_after:
        return None          # silent staleness = treat as suspended
    return m.odds
</code></pre>
<p>Two details matter here:</p>
<p>tradable_odds returns None for anything uncertain. Callers can't accidentally show a bad price because there is no bad price to get.
The staleness check handles the "silent suspension" case, where the provider simply stops sending updates.
Step 3: Handle reconnects without gaps</p>
<p>The scariest moment for a live odds feed is a dropped connection. During the gap, markets may have suspended, reopened, or settled, and you'd never know.</p>
<p>The safe pattern:</p>
<p>Mark everything as unverified the instant the socket drops.
Reconnect with exponential backoff.
Pull a REST snapshot to rebuild current state.
Resume applying live messages, using sequence numbers to skip anything the snapshot already covered.
python
import asyncio, json
import websockets</p>
<p>async def run_feed(book: MarketBook, ws_url: str, api_key: str, fetch_snapshot):
    backoff = 1
    while True:
        try:
            # NOTE: older versions of <code>websockets</code> call this param <code>extra_headers</code>
            async with websockets.connect(
                ws_url,
                additional_headers={"Authorization": f"Bearer {api_key}"},
                ping_interval=15,
            ) as ws:
                backoff = 1</p>
<pre><code>            # Anything we knew before the drop is no longer trustworthy
            for m in book.markets.values():
                if m.state is MarketState.OPEN:
                    m.state = MarketState.SUSPENDED

            # Rebuild truth from a snapshot, then continue streaming
            for msg in await fetch_snapshot():
                book.apply(msg)

            async for raw in ws:
                book.apply(json.loads(raw))

    except Exception as exc:
        print(f"feed error: {exc!r}, retrying in {backoff}s")

    await asyncio.sleep(backoff)
    backoff = min(backoff * 2, 30)
</code></pre>
<p>Look at the loop that downgrades OPEN to SUSPENDED on disconnect. That single block prevents the most dangerous class of bug: serving frozen "open" prices from a dead connection.</p>
<p>Step 4: Don't forget webhooks (and make them idempotent)</p>
<p>If you consume events through webhooks instead of streaming, the failure mode shifts. Providers retry on failure, so you will receive duplicates.</p>
<p>python
from fastapi import FastAPI, Request, HTTPException</p>
<p>app = FastAPI()
seen_event_ids: set[str] = set()      # use Redis/DB with TTL in production</p>
<p>@app.post("/webhook/odds")
async def odds_webhook(request: Request):
    payload = await request.json()</p>
<pre><code>event_id = payload.get("event_id")
if event_id is None:
    raise HTTPException(status_code=400, detail="missing event_id")

if event_id in seen_event_ids:
    return {"ok": True, "duplicate": True}   # ack, but don't reprocess
seen_event_ids.add(event_id)

book.apply(payload)                          # same state machine as above
return {"ok": True}
</code></pre>
<p>Rules of thumb for webhooks:</p>
<p>Acknowledge fast (return 2xx quickly) and process asynchronously.
Dedupe on an event ID.
Verify authenticity using whatever signing mechanism your provider supports.
Don't assume order. Use the same sequence-number guard.
Step 5: Make the UI honest</p>
<p>The backend can be perfect and the UI can still mislead. Good odds UIs follow a few rules:</p>
<p>Show suspension explicitly</p>
<p>Replace the price with a lock icon or "Suspended". Don't blank the cell, because a blank looks like a bug.</p>
<p>Disable, don't hide</p>
<p>Keep the selection visible but not clickable. Users should understand the market exists and will likely return.</p>
<p>Handle the bet slip</p>
<p>If a market suspends while a selection is in the slip:</p>
<p>Freeze the selection and show a warning
Never silently update the price and let the user confirm the old one
Offer "accept new odds" only after the market reopens
Animate price changes, not suspensions</p>
<p>A brief green/red flash on change is helpful. Flashing on every suspend/reopen cycle is noise.</p>
<p>A minimal React hook:</p>
<p>tsx
type Odds = { home: number; draw: number; away: number };</p>
<p>export function useTradableOdds(matchId: string, market: string) {
  const [odds, setOdds] = useState&lt;Odds | null&gt;(null);
  const [state, setState] = useState&lt;"open" | "suspended" | "closed" | "settled"&gt;("suspended");</p>
<p>  useEffect(() =&gt; {
    const unsub = oddsStore.subscribe(matchId, market, (m) =&gt; {
      setState(m.state);
      setOdds(m.state === "open" ? m.odds : null);   // never expose odds unless open
    });
    return unsub;
  }, [matchId, market]);</p>
<p>  return { odds, state, locked: state !== "open" };
}
Common mistakes (and how to avoid them)
Mistake	Why it hurts	Fix
Using a boolean available	Can't tell suspended from closed from settled	Use a real state machine
Deleting markets on suspend	Loses history, breaks open bet slips	Keep the row, change the state
Showing last-known price while suspended	Users bet on invalid odds	Return None unless open
Ignoring staleness	Silent feed death = frozen "live" prices	Add a TTL and enforce it
Trusting the socket after a drop	Missed suspensions during the gap	Downgrade + snapshot resync
No sequence handling	Old messages overwrite new state	Drop seq &lt;= last_seq
Treating reopen as the same price	Odds usually change after a suspension	Always re-read prices
Not storing transitions	Can't debug "odds vanished" tickets	Log every state change with timestamps
Turning suspensions into an advantage</p>
<p>Suspensions aren't only a problem to defend against. They're data.</p>
<p>Frequency and duration of suspensions correlate with match volatility. A match that suspends 14 times is a different product from one that suspends twice.
Price before vs. after a suspension shows how much the market moved on new information.
Line movement across the match powers charts, alerts, and trading signals.
Storing every transition lets you replay and backtest, and that's gold for analytics and ML teams.</p>
<p>Log each transition like this:</p>
<p>sql
CREATE TABLE market_events (
    id           BIGSERIAL PRIMARY KEY,
    match_id     TEXT        NOT NULL,
    market       TEXT        NOT NULL,
    from_state   TEXT,
    to_state     TEXT        NOT NULL,
    odds_before  JSONB,
    odds_after   JSONB,
    seq          BIGINT,
    received_at  TIMESTAMPTZ NOT NULL DEFAULT now()
);</p>
<p>CREATE INDEX idx_market_events_lookup
    ON market_events (match_id, market, received_at);
Why a normalized feed makes this dramatically easier</p>
<p>Every bookmaker signals suspension differently: one uses a status string, another nulls the odds, a third removes the market entirely. If you integrate bookmakers directly, you write and maintain a suspension parser per source, and each one breaks differently.</p>
<p>That's why many teams consume odds through a normalizing layer. Orbistats is built around that idea: one consistent schema for odds across 13 sports (football, basketball, American football, cricket, tennis, baseball, esports, combat sports, volleyball, handball, ice hockey, golf, and horse racing), so you write your state machine once.</p>
<p>How the pieces map to what you just built:</p>
<p>Markets and prices: The Odds API covers 1X2, moneyline, spreads, handicaps, totals, player props, futures, pre-match and live odds, and line movement, with bookmaker odds normalized into one schema.
Streaming: The WebSocket API gives you a persistent connection for live updates instead of repeated polling. That's the foundation of Step 3.
Push delivery: Webhooks deliver events to your server, which is the pattern from Step 4.
Match context: The Live Scores API supplies goals, cards, and substitutions, so you can explain why a market just suspended.
Schedules and results: The Sports Data API provides fixtures, standings, and settled results.
Backtesting: The Historical Sports Data API offers a multi-season archive (stated coverage back to 2010) including closing odds.
No-code display: Prefer not to build UI? Drop-in Widgets like Live Score and Odds Board can be embedded directly.</p>
<p>If you're building for professional use cases, the Sportsbooks &amp; Trading and Trading Desks pages describe the low-latency, line-movement-oriented setups in more detail.</p>
<p>Get hands-on in 10 minutes
Grab a free key at orbistats.com/signup.
Walk through the Quickstart. Auth is a Bearer token against <a href="https://api.orbistats.com/v1/">https://api.orbistats.com/v1/</a>.
Experiment without code in the Sandbox.
Use the API Reference to confirm exact field names for odds and status. The payloads in this post are illustrative.
python
import requests</p>
<p>API_KEY = "YOUR_API_KEY"
BASE = "<a href="https://api.orbistats.com/v1">https://api.orbistats.com/v1</a>"</p>
<p>resp = requests.get(
    f"{BASE}/football/matches/live",
    headers={"Authorization": f"Bearer {API_KEY}"},
    timeout=10,
)
resp.raise_for_status()
print(resp.json())</p>
<p>The free tier is documented at 150 requests/day, enough to prototype your state machine against real data. See pricing when you're ready to scale. New to the vocabulary? The Glossary and Guides cover odds formats, implied probability, and how WebSockets work.</p>
<p>Production checklist</p>
<p>Before shipping any live-odds feature, confirm:</p>
<p> Market state is an enum, not a boolean
 Unknown markets default to not tradable
 A single function gates everything shown as bettable
 Updates are idempotent and guarded by sequence numbers
 A staleness TTL is enforced
 Disconnects downgrade open markets and trigger snapshot resync
 Webhooks are deduplicated and acknowledged fast
 The UI shows suspended explicitly and protects the bet slip
 Every transition is logged with timestamps
 You've tested with simulated suspensions, including mid-tap
Key takeaways
A suspension is temporary and normal, not an error.
Model four states: open, suspended, closed, settled.
Fail closed: unknown or stale means not tradable.
Sequence numbers + idempotency protect you from duplicates and reordering.
Reconnects are the danger zone. Downgrade, snapshot, then resume.
The UI must never display a price it can't stand behind.
Log transitions. Suspension history is valuable data.
A normalized odds feed means one handler instead of one per bookmaker.
Final thought</p>
<p>The best odds experiences don't feel fast because they show prices constantly. They feel trustworthy because they stop showing a price the instant it stops being true. Get suspensions right, and users never notice them. Get them wrong, and it's the only thing they'll remember.</p>
<p>If this helped, share it with the teammate who's about to cache "last known odds" for 30 seconds. 😄</p>
]]></content:encoded></item><item><title><![CDATA[Fixture Mapping Hell: How to Match the Same Match Across Multiple Sports Data Providers]]></title><description><![CDATA[The 2 AM bug that started this
You integrate two sports data providers. One is cheap and fast, the other has better coverage. You ship. Then a user messages you:
"Why is Arsenal vs Chelsea showing up ]]></description><link>https://bettechmagnetics.hashnode.dev/fixture-mapping-hell-how-to-match-the-same-match-across-multiple-sports-data-providers</link><guid isPermaLink="true">https://bettechmagnetics.hashnode.dev/fixture-mapping-hell-how-to-match-the-same-match-across-multiple-sports-data-providers</guid><category><![CDATA[Sports API]]></category><category><![CDATA[API Design]]></category><category><![CDATA[Python]]></category><category><![CDATA[data-engineering]]></category><category><![CDATA[webdev]]></category><dc:creator><![CDATA[Vijay Choudhary]]></dc:creator><pubDate>Wed, 07 Oct 2026 16:14:11 GMT</pubDate><enclosure url="https://cdn.hashnode.com/uploads/covers/6a47ac8dec0b8a17b4dc915d/317ac7cb-3470-4ccc-8d6b-cd5310d9ac34.png" length="0" type="image/jpeg"/><content:encoded><![CDATA[<p>The 2 AM bug that started this</p>
<p>You integrate two sports data providers. One is cheap and fast, the other has better coverage. You ship. Then a user messages you:</p>
<p>"Why is Arsenal vs Chelsea showing up twice, once with odds and once with the live score?"</p>
<p>Welcome to fixture mapping hell.</p>
<p>Every provider has its own IDs, its own spelling of team names, its own idea of what "kickoff time" means, and its own way of handling postponements. Nothing lines up. And the moment you want to combine live scores from one source with odds from another, you need to answer a deceptively simple question:</p>
<p>Is fixture A-48213 from Provider A the same real-world match as fixture 9f3c-77 from Provider B?</p>
<p>This post covers why that's hard, how to solve it properly in Python, and when you should stop solving it yourself.</p>
<p>Why matching fixtures is harder than it looks</p>
<p>If every provider used the same IDs, this would be a join. They don't, so here is what you're actually up against.</p>
<ol>
<li>Team name chaos</li>
</ol>
<p>The same club shows up as:</p>
<p>Manchester City
Man City
Manchester City FC
Manchester City F.C.
Man. City</p>
<p>Add diacritics (Atlético Madrid vs Atletico Madrid), transliteration differences, and youth or women's team suffixes (U21, W), and exact string matching falls apart immediately.</p>
<ol>
<li>Kickoff time drift</li>
</ol>
<p>Providers disagree on timestamps more than you'd expect:</p>
<p>One stores UTC, another stores local time without a timezone.
One updates the time when a broadcaster moves the game, the other doesn't for hours.
Some give the scheduled time, others the actual start.</p>
<p>A ±5 minute drift is normal. A 3-hour drift usually means a timezone bug, or a different match entirely.</p>
<ol>
<li>Home/away flips</li>
</ol>
<p>Neutral-venue games (cup finals, international tournaments) are routinely listed as A vs B by one provider and B vs A by another. If you assume home is home, you'll miss the match, or worse, attach odds to the wrong side.</p>
<ol>
<li>Postponements and reschedules</li>
</ol>
<p>A fixture moves from Saturday to Wednesday. Provider A creates a new fixture ID. Provider B updates the existing one. Now your mapping table points at a dead record.</p>
<ol>
<li>Sport-specific weirdness</li>
</ol>
<p>The moment you leave football, the assumptions break:</p>
<p>Sport	What breaks
Baseball	Doubleheaders: same teams, same day, two games
Cricket	Multi-day Test matches have no single kickoff
Tennis	Players, not teams; doubles pairs; retirements
Golf	No home/away at all, just a field of players
Horse racing	Identified by track + race number, not "teams"
Combat sports	Fight cards with many bouts and late changes
Esports	Best-of-3 series vs individual maps</p>
<p>A mapper tuned for football will quietly produce garbage for the others.</p>
<p>The strategy: blocking, scoring, assignment</p>
<p>Don't compare every fixture to every other fixture. Treat it as a classic entity resolution problem with three stages:</p>
<p>Blocking. Narrow the candidate pool cheaply (same sport, same day ±1).
Scoring. Rate each candidate pair with a weighted similarity function.
Assignment. Enforce one-to-one matches so one fixture can't claim two partners.</p>
<p>Here's a working implementation.</p>
<p>Step 1: Normalize names
python
import re
import unicodedata</p>
<p>STOPWORDS = {"fc", "cf", "afc", "sc", "ac", "fk", "sk", "club", "the"}</p>
<p>ALIASES = {
    "man city": "manchester city",
    "man utd": "manchester united",
    "spurs": "tottenham hotspur",
    "inter": "internazionale",
}</p>
<p>def normalize_name(name: str) -&gt; str:
    # strip accents: "Atlético" -&gt; "Atletico"
    s = unicodedata.normalize("NFKD", name).encode("ascii", "ignore").decode()
    s = s.lower()
    s = re.sub(r"[^a-z0-9 ]", " ", s)          # drop punctuation
    tokens = [t for t in s.split() if t not in STOPWORDS]
    s = " ".join(tokens)
    return ALIASES.get(s, s)</p>
<p>An alias table feels crude, but in production it's the single highest-ROI piece of the whole system. Keep it in a database, not in code, so non-engineers can fix mappings.</p>
<p>Step 2: A common fixture shape</p>
<p>Before comparing anything, convert every provider's payload into one internal structure.</p>
<p>python
from dataclasses import dataclass
from datetime import datetime, timezone</p>
<p>@dataclass(frozen=True)
class Fixture:
    provider: str
    provider_id: str
    sport: str
    competition: str
    home: str
    away: str
    kickoff: datetime      # always timezone-aware, always UTC
    status: str = "scheduled"</p>
<pre><code>@property
def home_n(self): return normalize_name(self.home)
@property
def away_n(self): return normalize_name(self.away)
</code></pre>
<p>def to_utc(ts: str) -&gt; datetime:
    dt = datetime.fromisoformat(ts.replace("Z", "+00:00"))
    if dt.tzinfo is None:
        raise ValueError(f"Naive timestamp from provider: {ts}")
    return dt.astimezone(timezone.utc)</p>
<p>Notice that to_utc raises on naive timestamps instead of guessing. Silent timezone assumptions are the root cause of most mapping bugs. Fail loudly, and fix the adapter for that provider.</p>
<p>Step 3: Score a candidate pair
python
from rapidfuzz import fuzz</p>
<p>def score_pair(a: Fixture, b: Fixture, max_minutes: int = 180):
    gap = abs((a.kickoff - b.kickoff).total_seconds()) / 60
    if gap &gt; max_minutes:
        return 0.0, False</p>
<pre><code>time_score = 1 - gap / max_minutes

# straight orientation
straight = (
    fuzz.token_set_ratio(a.home_n, b.home_n) +
    fuzz.token_set_ratio(a.away_n, b.away_n)
) / 200

# swapped orientation (neutral venues)
swapped = (
    fuzz.token_set_ratio(a.home_n, b.away_n) +
    fuzz.token_set_ratio(a.away_n, b.home_n)
) / 200

is_swapped = swapped &gt; straight
name_score = max(straight, swapped)

return 0.7 * name_score + 0.3 * time_score, is_swapped
</code></pre>
<p>Names carry 70% of the weight and time carries 30%. A perfect name match 90 minutes off still passes. A perfect name match 3 hours off does not. Tune these weights against your own labeled data.</p>
<p>Step 4: Block, score, and assign one-to-one
python
from collections import defaultdict
from datetime import timedelta</p>
<p>def match_fixtures(base, other, threshold=0.82):
    # Blocking: index by (sport, date)
    index = defaultdict(list)
    for f in other:
        index[(f.sport, f.kickoff.date())].append(f)</p>
<pre><code>candidates = []
for a in base:
    for offset in (-1, 0, 1):                       # catch midnight-UTC edge cases
        day = a.kickoff.date() + timedelta(days=offset)
        for b in index.get((a.sport, day), []):
            s, swapped = score_pair(a, b)
            if s &gt;= threshold:
                candidates.append((s, swapped, a, b))

# Greedy one-to-one assignment, best scores first
candidates.sort(key=lambda c: c[0], reverse=True)
used_a, used_b, matches = set(), set(), []
for s, swapped, a, b in candidates:
    if a.provider_id in used_a or b.provider_id in used_b:
        continue
    used_a.add(a.provider_id)
    used_b.add(b.provider_id)
    matches.append({
        "a": a.provider_id,
        "b": b.provider_id,
        "confidence": round(s, 3),
        "home_away_swapped": swapped,
    })
return matches
</code></pre>
<p>The one-to-one rule is what handles baseball doubleheaders. Two games between the same teams on the same day are separated by kickoff time, and the greedy assignment makes sure game 1 can't steal game 2's partner.</p>
<p>Step 5: Quick sanity test
python
a = Fixture("prov_a", "A-48213", "football", "Premier League",
            "Manchester City FC", "Arsenal", to_utc("2026-10-15T19:00:00Z"))
b = Fixture("prov_b", "9f3c-77", "football", "EPL",
            "Man City", "Arsenal FC", to_utc("2026-10-15T19:05:00Z"))</p>
<p>print(match_fixtures([a], [b]))</p>
<h1>[{'a': 'A-48213', 'b': '9f3c-77', 'confidence': 0.99, 'home_away_swapped': False}]</h1>
<p>Persisting the mapping: the crosswalk table</p>
<p>Matching is useless if you recompute it on every request. Persist the result as a crosswalk keyed on your own canonical ID.</p>
<p>sql
CREATE TABLE canonical_fixture (
    canonical_id  UUID PRIMARY KEY,
    sport         TEXT NOT NULL,
    kickoff_utc   TIMESTAMPTZ NOT NULL,
    status        TEXT NOT NULL DEFAULT 'scheduled'
);</p>
<p>CREATE TABLE fixture_crosswalk (
    provider      TEXT NOT NULL,
    provider_id   TEXT NOT NULL,
    canonical_id  UUID NOT NULL REFERENCES canonical_fixture(canonical_id),
    confidence    REAL NOT NULL,
    swapped       BOOLEAN NOT NULL DEFAULT FALSE,
    verified_by   TEXT,                 -- 'auto' or reviewer name
    updated_at    TIMESTAMPTZ DEFAULT now(),
    PRIMARY KEY (provider, provider_id)
);</p>
<p>CREATE INDEX idx_crosswalk_canonical ON fixture_crosswalk (canonical_id);</p>
<p>A few rules that save you later:</p>
<p>Never overwrite silently. Keep confidence and verified_by so you can audit.
Route low-confidence matches (0.70–0.85) to a human review queue instead of auto-accepting them.
Store the swapped flag. If you attach odds, you must flip home and away prices when swapped = true.
Handling the nasty edge cases
Reschedules</p>
<p>When a fixture moves, don't create a new canonical record. Run a second pass for unmatched fixtures with a wider window (±7 days) and require a near-perfect name score (≥ 0.95). If it hits, update kickoff_utc on the existing canonical row.</p>
<p>Postponed and cancelled</p>
<p>Keep the canonical row, change status, and let consumers decide how to render it. Deleting rows breaks every downstream foreign key.</p>
<p>Individual sports</p>
<p>For tennis, golf, and combat sports, you aren't matching teams. You're matching participants inside an event. Add a participant-level crosswalk, and for tennis, normalize "Djokovic N." vs "Novak Djokovic" with a last-name-plus-initial rule before fuzzy matching.</p>
<p>Time-zone landmines</p>
<p>If two providers disagree by exactly 1, 2, 3, 5.5, or 8 hours, it's almost certainly a timezone bug rather than a data difference. Add an alert for gaps that match whole-hour offsets.</p>
<p>The honest cost of doing this yourself</p>
<p>Let's be real about what you're signing up for:</p>
<p>Task	Ongoing effort
Maintain alias tables for every league	Weekly
Review low-confidence matches	Daily at scale
Handle each provider's schema changes	Every release
Debug "why is this match missing odds?" tickets	Constantly
Repeat all of the above per sport	× 13</p>
<p>And that's before you add odds. If you also merge bookmaker feeds, you're building a parser per bookmaker on top of the fixture mapper.</p>
<p>Fixture mapping is an infrastructure problem, not a product feature. Your users don't care that you solved it. They only notice when you don't.</p>
<p>The alternative: one provider, one schema, one ID</p>
<p>The cleanest fix to fixture mapping hell is to not have a mapping problem at all.</p>
<p>If fixtures, live scores, statistics, and odds all come from a single normalized source, one match_id flows through everything. That's the approach Orbistats takes: a unified sports data and odds platform covering 13 sports (football, basketball, American football, cricket, tennis, baseball, esports, combat sports, volleyball, handball, ice hockey, golf, and horse racing) behind one consistent schema.</p>
<p>In practice, the same identifier appears across products:</p>
<p>http
GET <a href="https://api.orbistats.com/v1/football/matches/live">https://api.orbistats.com/v1/football/matches/live</a>
Authorization: Bearer YOUR_API_KEY
json
{
  "match_id": "match_50231",
  "status": "live",
  "minute": 72,
  "home": { "name": "Manchester City", "score": 2 },
  "away": { "name": "Arsenal", "score": 1 }
}</p>
<p>Here is how the pieces line up:</p>
<p>Schedules and results come from the Sports Data API (fixtures, results, standings, teams, competitions).
In-play events such as goals, cards, and substitutions come from the Live Scores API.
Pre-match and live prices come from the Odds API, where bookmaker odds are normalized into one schema, so you don't maintain a parser per bookmaker.
Team and player numbers (including xG) come from the Sports Statistics API.
Backtesting and ML datasets come from the Historical Sports Data API, which the site describes as a multi-season archive going back to 2010 for stated coverage.</p>
<p>And for delivery you can choose how data reaches you, instead of polling:</p>
<p>Persistent live streams over the WebSocket API
Push notifications on goals, cards, and status changes via Webhooks</p>
<p>When everything shares one schema, a goal event, the current score, and the live odds for the same match all hang off the same record. Your crosswalk table disappears, and with it the alias lists, the review queue, and the 2 AM pages.</p>
<p>Try it in 10 minutes</p>
<p>If you want to see whether a unified feed fits your stack before ripping out your mapper:</p>
<p>Create a free account and grab a key at orbistats.com/signup.
Follow the Quickstart. The auth flow is just a Bearer token against <a href="https://api.orbistats.com/v1/">https://api.orbistats.com/v1/</a>.
Poke at responses without writing code in the Sandbox.
Read the full Documentation and API Reference for fixtures, odds, lineups, events, players, and more.</p>
<p>A minimal first request in Python:</p>
<p>python
import requests</p>
<p>API_KEY = "YOUR_API_KEY"
BASE = "<a href="https://api.orbistats.com/v1">https://api.orbistats.com/v1</a>"</p>
<p>resp = requests.get(
    f"{BASE}/football/fixtures",
    headers={"Authorization": f"Bearer {API_KEY}"},
    timeout=10,
)
resp.raise_for_status()</p>
<p>for fx in resp.json().get("data", [])[:5]:
    print(fx)</p>
<p>Check the API Reference for the exact response envelope. The data key above is illustrative.</p>
<p>The free tier is documented at 150 requests/day, which is plenty for prototyping. Paid tiers are on the pricing page.</p>
<p>Quick decision guide: build or buy?</p>
<p>Build your own mapper if:</p>
<p>You must combine a proprietary internal feed with a public one
You need leagues no single provider covers
You have engineers to maintain alias tables long term</p>
<p>Use a unified provider if:</p>
<p>You're shipping a product, not a data pipeline
You cover multiple sports
You need odds, scores, and stats to line up without reconciliation
Your team's time is worth more than the API bill</p>
<p>If you're starting with one sport, look at the dedicated coverage pages for Football, Cricket, or Tennis to see what's available.</p>
<p>Key takeaways
Fixture mapping is entity resolution, not string comparison.
Use blocking → scoring → one-to-one assignment to stay fast and correct.
Always normalize to UTC and reject naive timestamps.
Handle home/away swaps explicitly, and flip odds when swapped.
Persist results in a crosswalk table with confidence and audit fields.
Send low-confidence matches to human review.
Every sport needs its own rules. Football logic will fail on baseball, tennis, and golf.
The cheapest mapping bug is the one you never have, so consider a single normalized source.</p>
<p>New to the terminology? The Glossary explains terms like xG and implied probability, and the Guides cover how odds APIs and WebSockets work.</p>
<p>Final thought</p>
<p>Fixture mapping is one of those problems that looks like a weekend project and quietly becomes a permanent team. Whether you build the matcher above or avoid the problem with a unified feed, the goal is the same: one match, one identity, everywhere in your app.</p>
<p>If this saved you from a 2 AM bug, share it with a teammate who's about to write if team_name == "Man City". 😄</p>
]]></content:encoded></item><item><title><![CDATA[From Event to Endpoint: Engineering a Real-Time Sports Data Architecture That Doesn't Lag]]></title><description><![CDATA[A goal is scored. Eleven seconds later, your app still shows 0 to 0. Meanwhile a user in another tab sees 1 to 0, and a third user sees the goal twice. Nothing crashed. Every service was up. The archi]]></description><link>https://bettechmagnetics.hashnode.dev/from-event-to-endpoint-engineering-a-real-time-sports-data-architecture-that-doesn-t-lag</link><guid isPermaLink="true">https://bettechmagnetics.hashnode.dev/from-event-to-endpoint-engineering-a-real-time-sports-data-architecture-that-doesn-t-lag</guid><category><![CDATA[Odds API]]></category><category><![CDATA[Sports API]]></category><category><![CDATA[websockets]]></category><category><![CDATA[latency]]></category><category><![CDATA[API development ]]></category><dc:creator><![CDATA[Vijay Choudhary]]></dc:creator><pubDate>Tue, 06 Oct 2026 16:49:17 GMT</pubDate><enclosure url="https://cdn.hashnode.com/uploads/covers/6a47ac8dec0b8a17b4dc915d/f3788f61-8882-43c4-bfa6-763fd0c616e3.png" length="0" type="image/jpeg"/><content:encoded><![CDATA[<p>A goal is scored. Eleven seconds later, your app still shows 0 to 0. Meanwhile a user in another tab sees 1 to 0, and a third user sees the goal twice. Nothing crashed. Every service was up. The architecture simply was not designed for the way live sports data actually behaves.</p>
<p>Real-time sports data is a deceptively hard problem. It looks like "receive an update, show an update", and it turns out to be a distributed systems problem with a very unforgiving audience: people who are watching the match on a TV and will notice instantly if your screen disagrees with it.</p>
<p>This post walks through a reference architecture for a real-time sports data pipeline, from the moment something happens on the pitch to the moment it lands in your user's app. It is written from the point of view of someone building on top of a sports data and odds API, and also useful if you are building the pipeline itself. We will cover where lag actually comes from, what each stage of the pipeline is responsible for, and the specific failure modes (duplicates, out of order events, dropped connections, slow consumers) that cause most real-world incidents. Along the way there is working code in TypeScript and Python that you can adapt.</p>
<p>A note on scope. What follows is a general reference design, not a description of the internals of any particular vendor, including ours. The goal is to give you a mental model and a set of patterns you can verify against whatever provider or stack you use.</p>
<p>Why sports data is harder than it looks</p>
<p>Most web data is relatively calm. A product price changes a few times a day. A sports match is the opposite, and four properties make it hard.</p>
<p>First, it is bursty. A quiet midweek afternoon might produce a trickle of updates. A packed Saturday with dozens of football matches, a full basketball slate and a cricket innings all running at once can produce orders of magnitude more events per second, and the peak is concentrated at exactly the moments users care about most: goals, wickets, red cards, final whistles.</p>
<p>Second, it is multi-source. A single fact, such as "the score is now 2 to 1", can arrive from more than one upstream feed at slightly different times and with slightly different shapes. Somebody has to decide which one wins.</p>
<p>Third, it is correctable. Sports data is revised after the fact. A goal is credited to a different player. A card is rescinded. A score is corrected after a video review. Your pipeline has to treat the stream as a sequence of statements that can be superseded, not as a log of immutable truth.</p>
<p>Fourth, it is order sensitive. "Goal" followed by "goal disallowed" is a completely different story from "goal disallowed" followed by "goal". If your pipeline can reorder events, your users will see wrong scores that look perfectly plausible.</p>
<p>Every pattern in this post exists to deal with one of those four properties.</p>
<p>The pipeline at a glance</p>
<p>Here is the whole journey, from the pitch to the screen:</p>
<p>text
   Source feeds / data collectors
            |
            v</p>
<ol>
<li>Ingest        (accept, validate, timestamp)
  |
  v</li>
<li>Normalize     (one schema, one set of IDs)
  |
  v</li>
<li>Deduplicate and order  (sequence per match)
  |
  v</li>
<li>State store   (current truth + recent history)
  |
  v</li>
<li>Fan-out       (push to subscribers)
  /      |       <br /> v       v        v
   REST     WebSocket   Webhooks
 \       |        /
  v      v       v</li>
<li>Client  (reconnect, resume, reconcile, render)</li>
</ol>
<p>Each box has one job, and most outages come from a box doing a second job it was not designed for. We will take them in order.</p>
<p>Stage 1: Ingest, and stamp everything on the way in</p>
<p>The ingest layer accepts raw updates from upstream sources. Its responsibilities are small and strict:</p>
<p>Validate that the message is well formed, and reject or quarantine what is not.
Attach an ingest timestamp from a trusted clock the moment the message arrives.
Record which source it came from.
Hand it off quickly, without doing heavy work inline.</p>
<p>The timestamp is the one people skip, and it is the most valuable thing you can add. If you stamp every message at the door, you gain the ability to measure the latency of every later stage by subtracting timestamps, and you gain a stable tiebreaker when two sources disagree. If you do not stamp at ingest, you will be guessing about where the seconds went during your first real incident.</p>
<p>Keep ingest dumb and fast. Anything slow here, such as an enrichment lookup or a database write on the hot path, turns a burst of events into a growing queue, and a growing queue is lag.</p>
<p>Stage 2: Normalize into one schema, once</p>
<p>Upstream feeds describe the same match in different ways. One calls the team "Man City", another "Manchester City FC", a third uses a numeric ID. One reports the minute as 72, another as "72:15", another as "72'+0". A consumer who has to handle every variant will eventually mishandle one.</p>
<p>Normalization maps every incoming shape into a single internal event schema with stable identifiers. This is the same idea that makes a unified sports data API valuable to developers: you integrate once against one shape instead of writing a parser per source. The principle holds on both sides of the API boundary, whether you are the provider or the consumer.</p>
<p>Here is a typed event schema and a normalizer in TypeScript. It is deliberately strict, because a loose schema is where bugs hide:</p>
<p>typescript
type EventType =
  | "goal"
  | "card"
  | "substitution"
  | "period_start"
  | "period_end"
  | "score_correction"
  | "event_voided";</p>
<p>interface NormalizedEvent {
  eventId: string;       // globally unique, stable across retries
  matchId: string;       // stable ID, never a display name
  type: EventType;
  minute: number | null; // parsed to a plain integer
  teamId: string | null;
  playerId: string | null;
  payload: Record&lt;string, unknown&gt;;
  source: string;
  ingestedAt: number;    // epoch ms, stamped at the door
}</p>
<p>interface RawFeedMessage {
  id?: string;
  match?: string | number;
  kind?: string;
  clock?: string | number;
  team?: string | number;
  player?: string | number;
  data?: Record&lt;string, unknown&gt;;
  source: string;
  receivedAt: number;
}</p>
<p>const KIND_MAP: Record&lt;string, EventType&gt; = {
  GOAL: "goal",
  YC: "card",
  RC: "card",
  SUB: "substitution",
  KO: "period_start",
  HT: "period_end",
  FT: "period_end",
  CORR: "score_correction",
  VOID: "event_voided",
};</p>
<p>// "72", "72:15", "72'+0" and 72 all become the integer 72.
function parseMinute(clock: string | number | undefined): number | null {
  if (clock === undefined || clock === null) return null;
  if (typeof clock === "number") return Math.floor(clock);
  const match = /^(\d+)/.exec(clock.trim());
  return match ? parseInt(match[1], 10) : null;
}</p>
<p>export function normalize(raw: RawFeedMessage): NormalizedEvent | null {
  const type = KIND_MAP[(raw.kind ?? "").toUpperCase()];
  if (!type || raw.match === undefined) return null; // quarantine upstream</p>
<p>  const matchId = String(raw.match);
  const minute = parseMinute(raw.clock);</p>
<p>  return {
    // Deterministic ID: the same fact from the same source always
    // produces the same eventId, which makes retries safe.
    eventId: raw.id ?? <code>${raw.source}:${matchId}:${raw.kind}:${minute}:${raw.team ?? ""}:${raw.player ?? ""}</code>,
    matchId,
    type,
    minute,
    teamId: raw.team !== undefined ? String(raw.team) : null,
    playerId: raw.player !== undefined ? String(raw.player) : null,
    payload: raw.data ?? {},
    source: raw.source,
    ingestedAt: raw.receivedAt,
  };
}</p>
<p>The detail worth stealing is the deterministic eventId. If the same fact arrives twice, from a retry or a second source, it produces the same ID, and every later stage can recognise it as a duplicate instead of treating it as a new goal.</p>
<p>Stage 3: Deduplicate and order, per match</p>
<p>This stage is where most "the score was wrong for a moment" bugs live.</p>
<p>Deduplication answers: have I already applied this event? Ordering answers: in what order should events for this match be applied? Both can be solved with one idea: give every match its own monotonically increasing sequence number, and process each match's events strictly in that order.</p>
<p>Why per match rather than global? Because a single global order would force unrelated matches to wait for each other, which adds exactly the lag you are trying to remove. Matches are independent, so you can process them in parallel, as long as events within one match stay ordered. In a log-based system this is usually done by partitioning on matchId, so that all events for a match land on the same partition and are consumed in order. Redis Streams is a good lightweight way to get an ordered, replayable per-key log if you do not want to run a full streaming platform.</p>
<p>Here is a small sequencer in Python that dedupes by event ID and assigns a per-match sequence number:</p>
<p>python
from collections import OrderedDict, defaultdict
from dataclasses import dataclass, field</p>
<p>@dataclass
class MatchLog:
    next_seq: int = 1
    # Remember recent event IDs only, so memory stays bounded.
    seen: "OrderedDict[str, None]" = field(default_factory=OrderedDict)
    max_seen: int = 5000</p>
<p>class Sequencer:
    def <strong>init</strong>(self):
        self.matches = defaultdict(MatchLog)</p>
<pre><code>def accept(self, event: dict):
    """Return the event with a per-match `seq`, or None if it is a duplicate."""
    log = self.matches[event["matchId"]]
    event_id = event["eventId"]

    if event_id in log.seen:
        return None  # duplicate: drop silently, never apply twice

    log.seen[event_id] = None
    if len(log.seen) &gt; log.max_seen:
        log.seen.popitem(last=False)  # forget the oldest ID

    sequenced = dict(event)
    sequenced["seq"] = log.next_seq
    log.next_seq += 1
    return sequenced
</code></pre>
<p>if <strong>name</strong> == "<strong>main</strong>":
    s = Sequencer()
    goal = {"matchId": "m1", "eventId": "a", "type": "goal"}
    print(s.accept(goal))   # seq 1
    print(s.accept(goal))   # None, duplicate
    print(s.accept({"matchId": "m1", "eventId": "b", "type": "card"}))  # seq 2</p>
<p>Two notes on real deployments. First, the sequencer must run as a single writer per match, otherwise two instances will hand out the same sequence number. Partitioning by matchId gives you that for free. Second, the seq you assign here is the backbone of everything that follows: it is what lets a client detect that it missed something, and what lets a server resume a dropped connection exactly where it left off.</p>
<p>Handling corrections is a design choice, not a code trick. The cleanest approach is to never mutate history: a score correction or a voided event is simply another event in the log with a higher seq, and the current state is derived by applying events in order. That way, every consumer can reach the same answer by replaying the same log.</p>
<p>Stage 4: Keep current state and recent history separately</p>
<p>Consumers want two different things, and one data structure cannot serve both well.</p>
<p>The first is current state: "what is the score right now?" This must be a fast key lookup, ideally in memory or in a low-latency store, because it is what REST requests and new subscribers hit.</p>
<p>The second is recent history: "what happened in the last few minutes, in order?" This is what a reconnecting client needs to catch up, and what a late-joining widget uses to build a timeline.</p>
<p>Keep both. Update current state as events are applied, and keep a bounded ring of recent events per match. The bounded part matters: an unbounded history on the hot path is a memory leak waiting for a long match. Archive older events to cheap storage instead, which is also what feeds a historical data API for backtesting and analytics.</p>
<p>A useful discipline is to make the state a pure function of the event log. If you can always rebuild current state by replaying the log, then a crashed node, a corrupted cache, or a bad deploy is an inconvenience, not a disaster.</p>
<p>Stage 5: Fan-out, and the three ways to deliver</p>
<p>After an event is applied, it has to reach subscribers. There are three standard delivery methods, and they are not interchangeable. Most mature integrations use more than one.</p>
<p>REST is request and response. It is simple, cacheable and universal, and it is the right tool for fetching state on demand: fixtures, standings, a match snapshot when a page loads. It is the wrong tool for staying live, because the client has to keep asking. If you poll once a second, your freshness is limited by the polling interval no matter how fast the server is.</p>
<p>WebSocket is a persistent, bidirectional connection. The server pushes each update the moment it is ready, so there is no polling term at all. This is the right tool for anything a person watches in real time, such as a match center or a moving odds board. The protocol is defined in RFC 6455, and the MDN WebSockets guide is a friendly introduction. A WebSocket sports API is the usual way to consume this style of feed.</p>
<p>Webhooks invert the relationship: the provider sends an HTTP request to your server when something happens. They are ideal for back end reactions, such as settling a market, sending a push notification or updating a database, where you want your server to be told rather than to keep a connection open. See how a webhooks API is typically structured, and note that webhook delivery has its own correctness rules, which we cover below.</p>
<p>A good mental model: REST gets you the state, WebSocket keeps you live, webhooks trigger your back end. A typical live product uses REST for the first paint, WebSocket for the live updates, and webhooks for server-side side effects.</p>
<p>A WebSocket fan-out server that supports resume</p>
<p>The feature that separates a toy from a production feed is resume. Connections drop constantly: phones switch networks, laptops sleep, load balancers recycle. If a client reconnects and the server just starts sending new events, the client has a hole in its timeline. If the server replays everything it missed, the client heals itself.</p>
<p>Here is a compact Node.js fan-out server using the ws library. It keeps a bounded recent log per match and lets a client say "I have everything up to sequence N, send me what I missed":</p>
<p>javascript
import { WebSocketServer } from "ws";</p>
<p>const wss = new WebSocketServer({ port: 8080 });</p>
<p>// matchId -&gt; { events: [...bounded ring], subscribers: Set }
const matches = new Map();
const MAX_HISTORY = 500;
const MAX_BUFFERED_BYTES = 1_000_000; // slow-consumer cutoff</p>
<p>function getMatch(matchId) {
  if (!matches.has(matchId)) {
    matches.set(matchId, { events: [], subscribers: new Set() });
  }
  return matches.get(matchId);
}</p>
<p>// Called by the pipeline for every sequenced event.
export function publish(event) {
  const match = getMatch(event.matchId);
  match.events.push(event);
  if (match.events.length &gt; MAX_HISTORY) match.events.shift();</p>
<p>  const frame = JSON.stringify({ op: "event", data: event });
  for (const socket of match.subscribers) {
    // Backpressure: never let one slow client hold memory hostage.
    if (socket.bufferedAmount &gt; MAX_BUFFERED_BYTES) {
      socket.close(1013, "slow consumer, please reconnect and resume");
      continue;
    }
    socket.send(frame);
  }
}</p>
<p>wss.on("connection", (socket) =&gt; {
  socket.on("message", (raw) =&gt; {
    let msg;
    try {
      msg = JSON.parse(raw.toString());
    } catch {
      return;
    }</p>
<pre><code>if (msg.op === "subscribe" &amp;&amp; typeof msg.matchId === "string") {
  const match = getMatch(msg.matchId);
  const lastSeq = Number.isInteger(msg.lastSeq) ? msg.lastSeq : 0;

  // Resume: replay everything the client has not seen yet.
  const oldestKept = match.events.length ? match.events[0].seq : null;
  if (oldestKept !== null &amp;&amp; lastSeq + 1 &lt; oldestKept) {
    // Gap is older than our buffer: tell the client to refetch a snapshot.
    socket.send(JSON.stringify({ op: "resync_required", matchId: msg.matchId }));
  } else {
    for (const event of match.events) {
      if (event.seq &gt; lastSeq) {
        socket.send(JSON.stringify({ op: "event", data: event }));
      }
    }
  }
  match.subscribers.add(socket);
  socket.on("close", () =&gt; match.subscribers.delete(socket));
}
</code></pre>
<p>  });
});</p>
<p>Three production details are baked into that small example. The bounded history stops memory growth. The resync_required message handles the case where a client was away so long that the replay buffer no longer covers the gap, which is the honest answer, rather than silently sending a partial timeline. And the bufferedAmount check is backpressure: a slow client that cannot keep up is disconnected and told to resume, instead of letting its queue grow until the server runs out of memory.</p>
<p>Stage 6: The client is part of the architecture</p>
<p>Many latency complaints are actually client bugs. A perfect server feeding a naive client still produces a laggy, wrong app. The client has four responsibilities: reconnect politely, resume from where it stopped, detect gaps, and reconcile with an authoritative snapshot when something looks wrong.</p>
<p>Reconnection must use exponential backoff with jitter. Without jitter, a server restart causes every client to retry at the same instants, producing a thundering herd that knocks the server over again. The AWS Builders' Library has an excellent write up on timeouts, retries and backoff with jitter that is worth reading in full.</p>
<p>Here is a browser client that reconnects with jitter, resumes from the last sequence it applied, ignores duplicates and old events, and falls back to a snapshot when it detects a gap:</p>
<p>javascript
class LiveMatchClient {
  constructor({ url, matchId, onState, fetchSnapshot }) {
    this.url = url;
    this.matchId = matchId;
    this.onState = onState;
    this.fetchSnapshot = fetchSnapshot; // () =&gt; Promise&lt;{ seq, state }&gt;
    this.lastSeq = 0;
    this.state = {};
    this.attempt = 0;
    this.closedByUser = false;
    this.connect();
  }</p>
<p>  connect() {
    this.socket = new WebSocket(this.url);</p>
<pre><code>this.socket.onopen = () =&gt; {
  this.attempt = 0;
  this.socket.send(
    JSON.stringify({ op: "subscribe", matchId: this.matchId, lastSeq: this.lastSeq })
  );
};

this.socket.onmessage = (message) =&gt; {
  const frame = JSON.parse(message.data);
  if (frame.op === "resync_required") return this.resync();
  if (frame.op === "event") this.apply(frame.data);
};

this.socket.onclose = () =&gt; {
  if (!this.closedByUser) this.scheduleReconnect();
};
</code></pre>
<p>  }</p>
<p>  apply(event) {
    if (event.seq &lt;= this.lastSeq) return; // duplicate or stale: ignore</p>
<pre><code>if (event.seq !== this.lastSeq + 1) {
  // We skipped one or more events. Do not guess: reload the truth.
  return this.resync();
}

this.state = reduce(this.state, event); // your pure reducer
this.lastSeq = event.seq;
this.onState(this.state);
</code></pre>
<p>  }</p>
<p>  async resync() {
    const snapshot = await this.fetchSnapshot();
    this.state = snapshot.state;
    this.lastSeq = snapshot.seq;
    this.onState(this.state);
  }</p>
<p>  scheduleReconnect() {
    // Full jitter: random delay between 0 and an exponentially growing cap.
    const cap = Math.min(30_000, 500 * 2 ** this.attempt);
    const delay = Math.random() * cap;
    this.attempt += 1;
    setTimeout(() =&gt; this.connect(), delay);
  }</p>
<p>  close() {
    this.closedByUser = true;
    this.socket.close();
  }
}</p>
<p>// Replace with your own pure function that applies one event to the state.
function reduce(state, event) {
  switch (event.type) {
    case "goal":
      return { ...state, lastGoal: event, score: event.payload.score ?? state.score };
    default:
      return state;
  }
}</p>
<p>The two rules in apply are the whole game. If an event is at or below the last applied sequence, ignore it. If it skips ahead, stop and reload a snapshot rather than rendering a state you cannot vouch for. A screen that is briefly "loading" is far better than a screen that is confidently wrong.</p>
<p>Webhooks: three rules that save you</p>
<p>Webhooks look simple, which is why they cause so many quiet bugs. If you receive them, follow three rules.</p>
<p>Verify the signature. Anyone who finds your endpoint can post to it, so the sender signs each payload and you verify before trusting it. GitHub's guide to validating webhook deliveries shows the standard HMAC pattern, and most providers follow a variation of it. Always compare signatures in constant time.</p>
<p>Acknowledge fast, process later. Return a 2xx immediately and push the real work to a queue. If your handler is slow, the sender will time out and retry, and now you are processing the same event twice while still being slow.</p>
<p>Be idempotent. Delivery is at least once, which means duplicates are normal, not exceptional. Key your processing on the event ID and make handling the same ID twice a no-op.</p>
<p>Here is a minimal Express handler that does all three:</p>
<p>javascript
import crypto from "node:crypto";
import express from "express";</p>
<p>const app = express();
const SECRET = process.env.WEBHOOK_SECRET;
const processed = new Set(); // use Redis or a database unique key in production</p>
<p>// Capture the raw body: the signature is computed over the exact bytes.
app.use(express.json({ verify: (req, _res, buf) =&gt; { req.rawBody = buf; } }));</p>
<p>app.post("/webhooks/sports", (req, res) =&gt; {
  const received = req.get("x-signature") ?? ""; // header name varies by provider
  const expected = crypto.createHmac("sha256", SECRET).update(req.rawBody).digest("hex");</p>
<p>  const a = Buffer.from(received);
  const b = Buffer.from(expected);
  if (a.length !== b.length || !crypto.timingSafeEqual(a, b)) {
    return res.status(401).end();
  }</p>
<p>  const event = req.body;
  if (processed.has(event.eventId)) {
    return res.status(200).end(); // already handled: acknowledge, do nothing
  }
  processed.add(event.eventId);</p>
<p>  res.status(200).end(); // acknowledge immediately
  queueMicrotask(() =&gt; handleEvent(event)); // real work happens off the request path
});</p>
<p>function handleEvent(event) {
  // Enqueue to a durable job queue in production.
  console.log("processing", event.type, event.matchId);
}</p>
<p>app.listen(3000);</p>
<p>The header name and signature format differ between providers, so check the documentation of the one you use. The shape, though, stays the same.</p>
<p>Where the milliseconds actually go</p>
<p>If you add a timestamp at every stage, you can build a latency budget for the whole pipeline, and it tends to be humbling. A rough, illustrative budget for one event might look like this:</p>
<p>Stage	What happens	Typical share
Capture at source	Event is recorded at the venue or by a collector	Outside your control
Ingest and validate	Accept, stamp, hand off	Small
Normalize and sequence	Map to schema, dedupe, assign seq	Small
State update and fan-out	Apply, publish to subscribers	Small to moderate
Network to the consumer	Distance and last mile	Often the largest
Client processing and render	Parse, reduce, paint	Small, but easy to ruin</p>
<p>The pattern is consistent: the stages you fully control are usually fast, and the stages you do not control, capture at the source and the network, dominate. That has a practical consequence. Past a point, shaving another millisecond from your server buys nothing. The wins come from fewer hops, closer regions, push instead of poll, and a client that does not stall. If you want the measurement side of this, our companion post on why latency claims need a definition shows how to measure end to end with real scripts.</p>
<p>Observability: measure the lag your users feel</p>
<p>You cannot fix lag you cannot see. Instrument four things.</p>
<p>Stage latency. Track the time between consecutive timestamps (ingest to sequence, sequence to publish) as histograms, and look at p95 and p99, never only the average.</p>
<p>Event age at delivery. For every message sent, record how old it was at the moment it left. This is the number that tells you whether you are actually fresh, as opposed to merely responsive.</p>
<p>Queue depth and consumer lag. A queue that grows during a burst is the earliest warning of trouble, minutes before users notice.</p>
<p>Client-side gap and resync rate. If clients are frequently detecting gaps or requesting snapshots, something upstream is dropping or reordering events, even if every server dashboard looks green.</p>
<p>Alert on these, and especially on event age at the 99th percentile during peak load. Peak is where the pipeline earns its keep, and it is exactly where a quiet-hours benchmark tells you nothing.</p>
<p>Failure modes, and what to do about each</p>
<p>A short field guide to the incidents you will actually have:</p>
<p>Duplicate events. Cause: retries, multiple sources. Defence: deterministic event IDs and idempotent apply.
Out of order events. Cause: parallel processing, network reordering. Defence: per-match sequence numbers and a single writer per match.
Dropped connections. Cause: mobile networks, load balancer recycling. Defence: resume from last sequence with a bounded replay buffer.
Slow consumers. Cause: congested clients. Defence: backpressure limits and forced reconnect with resume.
Thundering herd on reconnect. Cause: synchronized retries after an outage. Defence: exponential backoff with full jitter.
Corrections and voided events. Cause: video review, data fixes. Defence: model them as new events, never mutate history.
State drift. Cause: a missed event nobody noticed. Defence: periodic reconciliation against an authoritative snapshot.
Traffic bursts. Cause: many matches ending together. Defence: partition by match, scale consumers horizontally, shed non-critical work first.
Choosing the right delivery method for your product</p>
<p>Do not pick a transport because it sounds modern. Pick it because of what your product does.</p>
<p>A match center or live scoreboard on a website wants a WebSocket for updates, plus REST for the first paint and for the resync snapshot. If you would rather not build the UI yourself, ready-made widgets can handle the front end.</p>
<p>A trading or odds product cares about event-to-delivery time and jitter, and should consume a stream. The odds API pattern of normalized markets over a persistent connection exists for exactly this reason.</p>
<p>A notification or settlement system is a natural fit for webhooks, with idempotent handlers and a durable queue behind them.</p>
<p>A model or analytics pipeline mostly needs completeness and history, and benefits from bulk historical pulls and scheduled REST calls rather than a live socket.</p>
<p>A mobile app should assume the connection will drop constantly, and should lean on resume and snapshots rather than hoping for a perfect link.</p>
<p>Putting it into practice</p>
<p>You can try the consumer side of everything above without building the server half. The documentation covers authentication and the base URL, the quickstart gets you to a first request, and the sandbox lets you experiment before you commit. The live scores API and the football coverage page are a good place to start if your first project is a match center, and the pricing page lists the plan limits so you can size a prototype realistically.</p>
<p>Orbistats covers 13 sports, from football, basketball and cricket to tennis, esports and horse racing, behind one consistent schema, with a free tier you can start on today at orbistats.com. Because the schema is shared across sports, the client code in this post works with minimal changes whether the event is a goal, a wicket or a set point.</p>
<p>The short version</p>
<p>A real-time sports data architecture that does not lag is built from boring, well understood pieces assembled with discipline. Stamp at the door. Normalize once, into one schema with stable IDs. Give every match its own ordered sequence and process it with a single writer. Keep current state and recent history separately, and derive state from the log. Push instead of poll. Make resume a first-class feature on the server, and gap detection plus snapshot recovery a first-class feature on the client. Make webhooks signed, fast to acknowledge and idempotent. Then measure event age at the 99th percentile on your busiest day, because that is the number your users actually experience.</p>
<p>None of these ideas is exotic. The difference between an app that feels instant and one that feels broken is almost never a clever algorithm. It is whether someone remembered to handle the duplicate, the reorder, the dropped connection and the slow consumer before the final whistle.</p>
]]></content:encoded></item><item><title><![CDATA[Why "Sub-50ms" Claims Are Mostly Meaningless Without Knowing Which Latency They Measure]]></title><description><![CDATA[Every API landing page has a number on it. "Sub-50ms." "Real-time." "Lightning fast." It sits next to a green dot and a globe animation, and it is supposed to settle the argument before you have writt]]></description><link>https://bettechmagnetics.hashnode.dev/why-sub-50ms-claims-are-mostly-meaningless-without-knowing-which-latency-they-measure</link><guid isPermaLink="true">https://bettechmagnetics.hashnode.dev/why-sub-50ms-claims-are-mostly-meaningless-without-knowing-which-latency-they-measure</guid><category><![CDATA[Odds API]]></category><category><![CDATA[Sports API]]></category><category><![CDATA[websockets]]></category><category><![CDATA[latency]]></category><category><![CDATA[API development ]]></category><dc:creator><![CDATA[Vijay Choudhary]]></dc:creator><pubDate>Tue, 06 Oct 2026 16:46:35 GMT</pubDate><enclosure url="https://cdn.hashnode.com/uploads/covers/6a47ac8dec0b8a17b4dc915d/27ee13e2-787a-4dab-9307-a2edd6de8e83.png" length="0" type="image/jpeg"/><content:encoded><![CDATA[<p>Every API landing page has a number on it. "Sub-50ms." "Real-time." "Lightning fast." It sits next to a green dot and a globe animation, and it is supposed to settle the argument before you have written a single line of code.</p>
<p>It does not settle anything, because "50ms" is not a measurement. It is the answer to a question the page forgot to print. Fifty milliseconds of what? Measured from where? At which percentile? On a warm connection or a cold one? Over REST or over a socket?</p>
<p>If you build anything live, whether that is a score ticker, a betting model, a fantasy app or a trading screen, the difference between those answers is the difference between a product that feels instant and one that feels broken. This post breaks the number apart. We run a sports data and odds platform, and our own sub-50ms live feed claim is exactly the kind of number this article is about, so treat everything below as a way to hold any latency claim accountable, ours included.</p>
<p>By the end you will know the five latencies that get mixed up under one label, the physics that sets a hard floor on all of them, why percentiles matter more than averages, why polling quietly destroys a fast API, and how to measure all of it with code you can run today.</p>
<p>There is no such thing as "the" latency</p>
<p>When a provider says "sub-50ms", they usually mean one of these five things. Each is real. Each is measured at a different point. Each can differ from the next by an order of magnitude.</p>
<p>Server processing time. How long the application code took between receiving a request and producing a response. This is the number that looks best, because it excludes the network, the TLS handshake, the queue in front of the server and everything that happens before the request reaches the code.
Time to first byte at the edge. How long until the first byte of the response arrives at a client, measured from a location the provider chose, usually close to the servers.
Full request and response, as seen by the caller. The complete round trip from your machine, including DNS, connection setup, transfer and your own client overhead.
Event to delivery. For a push feed, the time between a change being known inside the provider's system and the message arriving on your socket. This is the honest number for streaming, and the hardest to measure from the outside.
Real world to screen. The time between something happening on the pitch and your user seeing it. A goal is scored. How long until the pixel changes in your app? This is the only number your user experiences, and almost no landing page quotes it.</p>
<p>A claim of "sub-50ms" is credible for number 1 on almost any decent stack. It is genuinely hard for number 4. For number 5 it is close to impossible without describing exactly which sources, which sport and which delivery method are involved, because the data has to be captured at a venue before any API sees it.</p>
<p>None of this means a provider is lying. It means the number is incomplete. Your job as the engineer is to complete it.</p>
<p>Follow one goal from the pitch to your user</p>
<p>To see why the label matters, follow a single event through the pipeline. A goal is scored in a football match. What has to happen before your user sees it?</p>
<p>Stage	Who owns it	Usually inside the "sub-50ms" claim?
Event happens and is captured at the source	Data collectors, official feeds, venue systems	No
Provider ingests and validates the update	Provider	Sometimes
Provider normalizes it into a consistent schema	Provider	Sometimes
Provider fans it out to subscribers	Provider	Usually yes
Network travel to your server	The internet, and your region	No
Your parsing and business logic	You	No
Your push to the browser or the phone	You, plus the last mile	No</p>
<p>Look at the last column. The stages inside the claim are the ones the provider fully controls. The stages outside it are where most of the wall clock time actually goes. That is not deceptive, it is just scope. But if you pick a provider because of a number that covers two rows of this table, and your product feels slow, the number was never going to protect you.</p>
<p>The physics floor nobody can negotiate with</p>
<p>Before we talk about software, there is a hard limit. Light in optical fiber travels at roughly two thirds of its speed in a vacuum, about 200,000 kilometers per second. That works out to roughly 5 milliseconds of one way travel per 1,000 kilometers, or about 10 milliseconds of round trip per 1,000 kilometers, even if every router and every server in the world were infinitely fast.</p>
<p>Real routes are longer than the straight line, and real equipment adds delay on top. But even this idealized floor is useful, because it tells you what is possible:</p>
<p>If a client and a server are 5,000 km apart, the round trip cannot be under about 50 ms. Not with better code, not with a bigger server.
A "sub-50ms response" measured from inside the same region as the server tells you nothing about what a developer on another continent will see.
If your users are far from the provider, the only fixes are moving closer, moving the work to the edge, or switching from request response to push so you pay the distance once rather than on every poll.</p>
<p>So the first question to ask any provider is simply: measured from where?</p>
<p>Cold connections are expensive, warm connections are cheap</p>
<p>A fast server cannot hide the cost of setting up a connection. Over HTTPS, a brand new connection typically needs a TCP handshake, then a TLS handshake, then the request itself. With TLS 1.3 that is about three round trips before the first byte of the response arrives. At a 20 ms round trip that is around 60 ms before the server has done anything. At 80 ms it is 240 ms.</p>
<p>This is why the same endpoint can look wildly different in two benchmarks. A script that opens a fresh connection per request measures setup cost. A client that reuses one keep-alive connection measures mostly the request itself.</p>
<p>You can see the breakdown of a single request with curl:</p>
<p>bash
curl -s -o /dev/null <br />  -H "Authorization: Bearer $ORBISTATS_API_KEY" <br />  -w "dns:%{time_namelookup}s  tcp:%{time_connect}s  tls:%{time_appconnect}s  ttfb:%{time_starttransfer}s  total:%{time_total}s\n" <br />  <a href="https://api.orbistats.com/v1/football/matches/live">https://api.orbistats.com/v1/football/matches/live</a></p>
<p>Each value is cumulative from the start of the request, so the cost of each phase is the difference between neighbouring numbers. Run it five times and then run it again after the first call. You will usually see the first request carry the dns, tcp and tls cost, and later requests on a reused connection skip most of it, as long as your client keeps the connection alive.</p>
<p>A related trick: some APIs return a Server-Timing header that reports how long the server itself spent. If a provider exposes it, you can subtract it from your measured time to isolate network and queueing. The MDN reference for the Server-Timing header explains the format.</p>
<p>Averages lie, percentiles tell the truth</p>
<p>Suppose an API answers 99 requests in 20 ms and one request in 2,000 ms. The average is about 40 ms. A marketing page can honestly write "sub-50ms average". And one in a hundred of your users just waited two seconds.</p>
<p>That is why latency should always be described with percentiles:</p>
<p>p50 (the median): what a typical request feels like.
p95: what one request in twenty feels like.
p99: what one request in a hundred feels like.
max: the worst case you saw.</p>
<p>This matters even more than it looks, because real applications do not make one request. A single screen might make a dozen calls, and the page is as slow as the slowest of them. The more calls you fan out, the more often you hit someone's p99. Google's paper The Tail at Scale is the classic explanation of why rare slow responses dominate the experience of systems with many parts, and the monitoring chapter of the Google SRE book makes the same case for always looking at distributions rather than means.</p>
<p>So the second question to ask is: at which percentile? "Sub-50ms at p50" and "sub-50ms at p99" are completely different products.</p>
<p>Measure it yourself, and measure it honestly</p>
<p>Here is a small Python script that fires requests on a fixed schedule over one reused connection and reports percentiles. It deliberately measures each request from the moment it was supposed to start, not from the moment your code got around to sending it. That detail is called avoiding coordinated omission. A naive loop that waits for each response before sending the next one will quietly slow itself down during a stall and under-report exactly the slow moments you care about.</p>
<p>python
import asyncio
import os
import time</p>
<p>import httpx</p>
<p>URL = "<a href="https://api.orbistats.com/v1/football/matches/live">https://api.orbistats.com/v1/football/matches/live</a>"
HEADERS = {"Authorization": f"Bearer {os.environ['ORBISTATS_API_KEY']}"}</p>
<p>def percentile(values, p):
    ordered = sorted(values)
    index = min(len(ordered) - 1, round((p / 100) * (len(ordered) - 1)))
    return ordered[index]</p>
<p>async def timed_request(client, intended_start, results):
    # Wait until the moment this request was SCHEDULED to start.
    await asyncio.sleep(max(0, intended_start - time.perf_counter()))
    try:
        response = await client.get(URL, headers=HEADERS)
        response.raise_for_status()
        # Latency is measured from the intended start, so any delay
        # caused by a stalled loop or a slow earlier request still counts.
        results.append((time.perf_counter() - intended_start) * 1000)
    except httpx.HTTPError as error:
        print("request failed:", error)</p>
<p>async def main(count=50, interval_seconds=1.0):
    results = []
    async with httpx.AsyncClient(timeout=5.0) as client:
        # One warm-up call so the TCP and TLS cost is not mixed into the numbers.
        await client.get(URL, headers=HEADERS)
        start = time.perf_counter() + 0.5
        await asyncio.gather(
            *(
                timed_request(client, start + i * interval_seconds, results)
                for i in range(count)
            )
        )</p>
<pre><code>if not results:
    print("no successful requests")
    return

print(f"requests: {len(results)}")
print(f"p50: {percentile(results, 50):.1f} ms")
print(f"p95: {percentile(results, 95):.1f} ms")
print(f"p99: {percentile(results, 99):.1f} ms")
print(f"max: {max(results):.1f} ms")
</code></pre>
<p>asyncio.run(main())</p>
<p>A few practical notes. Every call counts against your plan's request quota, so keep the count modest on a free key and check the current limits on the pricing page. Run the script from the region where your production code will actually live, not from your laptop on hotel Wi-Fi. And run it at different times of day, because a single quiet afternoon proves very little. The sandbox is a good place to rehearse the first run before you point the script at anything that matters.</p>
<p>Notice what this script measures: number 3 from the list above, the full request and response as the caller sees it, on a warm connection. It says nothing about freshness. An API can answer in 15 ms with data that is already three seconds old. Response time and data age are two separate things, and only one of them is in the headline.</p>
<p>Polling quietly destroys a fast API</p>
<p>This is the part that surprises people most. Imagine an API with a perfect 10 ms response time, and an app that polls it once per second to show live scores.</p>
<p>When a goal is scored, it lands at a random moment inside the polling window. On average your app learns about it half an interval later, so about 500 ms. In the worst case it learns about it a full interval plus the response time later, so just over a second. The 10 ms response time barely matters, because the dominant term is the polling interval, not the speed of the server.</p>
<p>The rule of thumb:</p>
<p>text
worst-case staleness (polling) = polling interval + response time
average staleness (polling)    = polling interval / 2 + response time</p>
<p>You can shrink the interval, but then you multiply your request volume, burn through your quota and add load to both sides, and you still never get below the interval. The only way out is to stop asking and let the provider tell you. There are two standard ways to do that:</p>
<p>A persistent connection. A WebSocket feed keeps one connection open and the provider pushes each update the moment it is ready, so the polling term disappears entirely. The protocol is described in the MDN WebSockets guide.
A callback. Webhooks let the provider send an HTTP request to your server when something happens, which is a good fit when you want to trigger back end logic rather than keep a UI continuously fresh.</p>
<p>For anything that has to feel live, such as the live scores in a match center or a moving line in an odds board, push beats poll almost every time. The headline number for a REST endpoint and the headline number for a stream are not even the same kind of number, which is why you should never compare them directly.</p>
<p>How to measure a push feed</p>
<p>Measuring push is harder, because there is no request to time. What you want is event to delivery: how long after the provider produced a message did it arrive on your side?</p>
<p>If the provider includes a server timestamp in each message, you can compute the difference between your receive time and that timestamp. The catch is clocks. Your clock and the provider's clock are never perfectly aligned, so the absolute difference includes whatever offset exists between them. Two honest ways to deal with this:</p>
<p>Run your test machine with a well-synchronised clock, using NTP or chrony, and treat the result as an estimate with a known error margin.
Focus on the spread rather than the absolute value. A constant clock offset shifts every sample by the same amount, so the gap between p99 and p50 is unaffected by it. That gap is your jitter, and it is often the number that decides whether a feed is usable.</p>
<p>Here is a minimal Node.js client that does exactly that. The field name emitted_at is illustrative; replace it with whatever timestamp field the payload actually carries, and copy the socket address from the documentation.</p>
<p>javascript
import WebSocket from "ws";</p>
<p>const WS_URL = process.env.ORBISTATS_WS_URL; // copy the address from the docs
const API_KEY = process.env.ORBISTATS_API_KEY;
const RUN_MINUTES = 5;</p>
<p>const deltas = [];</p>
<p>const ws = new WebSocket(WS_URL, {
  headers: { Authorization: <code>Bearer ${API_KEY}</code> },
});</p>
<p>ws.on("message", (raw) =&gt; {
  const receivedAt = Date.now();
  let message;
  try {
    message = JSON.parse(raw.toString());
  } catch {
    return; // ignore anything that is not JSON
  }
  if (typeof message.emitted_at !== "number") return;
  deltas.push(receivedAt - message.emitted_at);
});</p>
<p>ws.on("error", (error) =&gt; console.error("socket error:", error.message));</p>
<p>function percentile(values, p) {
  const sorted = [...values].sort((a, b) =&gt; a - b);
  const index = Math.min(sorted.length - 1, Math.round((p / 100) * (sorted.length - 1)));
  return sorted[index];
}</p>
<p>setTimeout(() =&gt; {
  if (deltas.length === 0) {
    console.log("no timestamped messages received");
  } else {
    const p50 = percentile(deltas, 50);
    const p95 = percentile(deltas, 95);
    const p99 = percentile(deltas, 99);
    console.log({
      messages: deltas.length,
      p50_ms: p50,
      p95_ms: p95,
      p99_ms: p99,
      // Unaffected by a constant clock offset between you and the provider.
      jitter_p99_minus_p50_ms: p99 - p50,
    });
  }
  ws.close();
}, RUN_MINUTES * 60 * 1000);</p>
<p>Run it during a busy period, such as a full football match day or a packed evening of fixtures, not at 4 a.m. on a Tuesday. Light load flatters every system. If the spread between p50 and p99 grows sharply when many events fire at once, you have learned something no landing page will ever tell you.</p>
<p>Which latency does your product actually need?</p>
<p>The right target depends entirely on what you are building, and chasing a smaller number than you need is its own form of waste.</p>
<p>A trading or odds desk reacting to line movement lives and dies on event to delivery and on jitter. Here the difference between 50 ms and 500 ms is real money, and you care about the p99. This is the audience for odds data delivered over a stream.
A media site with a live match center does not need milliseconds. Users cannot perceive a difference of a few hundred milliseconds in a score update, but they absolutely notice a feed that stalls for ten seconds. Reliability and tail latency matter more than the median.
A fantasy or analytics product mostly cares about correctness and completeness, with freshness measured in seconds or minutes.
A mobile app on a mid range phone and a congested network will add far more latency in the last mile than any provider can remove.</p>
<p>Write down the number your product genuinely needs, then work backwards through the table from earlier, assigning a share of the budget to each stage. If the provider only owns two stages and you are asking them to cover the whole thing, that mismatch is worth knowing before you sign anything.</p>
<p>A checklist to use on any latency claim</p>
<p>Print this out, paste it into your vendor evaluation doc, and ask it of every provider, including us.</p>
<p>Which of the five latencies does the number describe: server time, edge time to first byte, full round trip, event to delivery, or real world to screen?
Measured from where, and from which region and network?
At which percentile: p50, p95, p99?
Over what time window, and under what load? Quiet hours or peak match day?
Cold connection or warm keep-alive connection?
REST, WebSocket, or webhook, and is the number the same across all three?
Does it include TLS and DNS, or only application time?
Does the number reflect data age as well as response speed?
Is it per request, or per message on a stream?
Is there a public status page and a history of the real numbers, rather than a single snapshot?
Can you test it on a free tier before committing?
Do the published numbers match what you measure from your own servers?</p>
<p>Questions 10 to 12 are the important ones, because they turn a claim into evidence. A provider that publishes a status page, lets you try the real API without a sales call, and invites you to measure it yourself is telling you something that no adjective can.</p>
<p>Going further</p>
<p>If you want the longer, end to end treatment of how a low latency live feed is put together, the research section has deeper write ups, including a piece on what sub-50ms actually requires from end to end. And if you would rather test than read, Orbistats offers a free tier across 13 sports, so you can point the two scripts above at a real feed this afternoon and see which of the five latencies you are actually getting.</p>
<p>The takeaway is simple. A latency number without a definition is a slogan. Ask which clock, from where, at what percentile, over which transport, and then measure it yourself. The number you end up with will be less impressive than the one on the landing page, and far more useful.</p>
]]></content:encoded></item><item><title><![CDATA[Designing a Low-Latency Sports Data Pipeline: Caching, WebSockets and Webhooks Together]]></title><description><![CDATA[Most tutorials treat REST, WebSockets and webhooks as competitors. "Which one should I use?" is the wrong question.
In a real sports product, each one fails in a different way, and each one is good at]]></description><link>https://bettechmagnetics.hashnode.dev/designing-a-low-latency-sports-data-pipeline-caching-websockets-and-webhooks-together</link><guid isPermaLink="true">https://bettechmagnetics.hashnode.dev/designing-a-low-latency-sports-data-pipeline-caching-websockets-and-webhooks-together</guid><category><![CDATA[System Design]]></category><category><![CDATA[websockets]]></category><category><![CDATA[APIs]]></category><category><![CDATA[Redis]]></category><category><![CDATA[Backend Development]]></category><dc:creator><![CDATA[Vijay Choudhary]]></dc:creator><pubDate>Tue, 06 Oct 2026 16:38:56 GMT</pubDate><enclosure url="https://cdn.hashnode.com/uploads/covers/6a47ac8dec0b8a17b4dc915d/48067134-d830-431b-ae59-a53cedf215a2.png" length="0" type="image/jpeg"/><content:encoded><![CDATA[<p>Most tutorials treat REST, WebSockets and webhooks as competitors. "Which one should I use?" is the wrong question.</p>
<p>In a real sports product, each one fails in a different way, and each one is good at a different job. A pipeline that depends on just one of them is only as reliable as that one's worst day. A pipeline that combines all three, with a single place where the data is reconciled, is much harder to break.</p>
<p>This article designs that pipeline end to end: ingestion, a version-guarded state store, a multi-tier cache, a fan-out gateway, a reconciliation loop, and a chaos test that proves the whole thing tolerates duplicates, reordering and dropped messages. Code is TypeScript on Node with Redis. The provider-specific parts use Orbistats, which exposes REST, WebSocket and webhook delivery behind one normalized schema across 13 sports, but the architecture applies to any provider.</p>
<p>Honesty note: I don't invent benchmark numbers here. Field names like version and anything marked "placeholder" must be mapped to your provider's real payloads before you ship. The verification list is at the end.</p>
<p>Table of contents
Three channels, three jobs
Design principles
The architecture at a glance
The data model and the version guard
Channel 1: WebSocket ingest
Channel 2: Webhooks through a durable stream
Channel 3: REST for snapshots, reconciliation and cold data
The cache hierarchy
The gateway: L1 cache, snapshot-on-connect, fan-out
Failure modes and the degradation ladder
Tuning per sport
Observability: the metrics that matter
Chaos testing the reducer
Capacity, cost and limits
Checklist and next steps</p>
<ol>
<li>Three channels, three jobs
Channel	Strength	Weakness	Best job
WebSocket	Lowest delay, continuous stream	Silent gaps when the connection drops	Primary live feed
Webhooks	Provider retries, pushes to you without a held connection	Duplicates, out-of-order retries, needs a public endpoint	Durable event notifications and side effects
REST	Simple, cacheable, gives you current truth on demand	Polling adds interval-sized staleness and burns request budget	Snapshots, reconciliation, fixtures, history</li>
</ol>
<p>Their weaknesses barely overlap. When the socket drops, webhooks and REST keep working. When your webhook endpoint is down, the socket still streams. When both push channels misbehave, REST can repair state. That complementarity is the entire point.</p>
<p>The product pages describe each channel from the provider side: Live Scores API, WebSocket API and Webhooks API. For pre-match and live prices, the Odds API uses the same delivery options, which matters if you serve trading-style users who are very sensitive to staleness.</p>
<ol>
<li>Design principles
One reducer, many inputs. Every channel funnels into the same function that decides whether an update is new. Channels never write state directly.
Idempotent and order-safe. Applying the same update twice, or an old update after a newer one, must be a no-op.
Snapshot plus delta. Anyone who connects, reconnects or restarts gets current state first, then live changes.
Push-invalidated caches. For live data, don't guess a TTL. Update caches the moment state changes. Use TTLs only for data that changes slowly.
Assume every channel lies sometimes. A periodic reconciliation pass compares your state against the source of truth and measures how often the push channels missed something.
Degrade, don't die. Define in advance what the product shows when each layer fails.</li>
<li>The architecture at a glance
text
 Provider
 ├─ WebSocket ─────► ingest-ws ──────┐
 ├─ Webhooks ──► receiver ► Stream ► worker ─┤
 └─ REST ──────► reconciler ─────────┤
                              ▼
                    ┌──────────────────┐
                    │ applyIfNewer()   │  ← the ONLY writer
                    │ (Lua, atomic)    │
                    └────────┬─────────┘
                             │ HSET + PUBLISH
                    ┌────────▼─────────┐
                    │  Redis (L2)      │  state hashes + pub/sub
                    └────────┬─────────┘
             ┌───────────────┼───────────────┐
             ▼               ▼               ▼
       Gateway A        Gateway B        Gateway C
       (L1 cache)       (L1 cache)       (L1 cache)
             │               │               │
          clients         clients         clients</li>
</ol>
<p>Everything funnels through one atomic function. That is what makes combining three unreliable channels safe.</p>
<ol>
<li>The data model and the version guard</li>
</ol>
<p>Each match has a single state object with a monotonically increasing version. Your provider should supply something usable here (a sequence number, a revision counter, or at worst an event timestamp). Check the API reference for what's actually available. Ordering by arrival time is the weakest fallback, because the three channels arrive at different speeds.</p>
<p>ts
// types.ts
export interface MatchState {
  matchId: string;
  sport: string;               // "football", "basketball", ...
  version: number;             // placeholder: map to the provider's sequence/revision
  status: "scheduled" | "live" | "finished";
  minute?: number;
  home: { name: string; score: number };
  away: { name: string; score: number };
}</p>
<p>Here is the reducer in its purest form. It has no I/O, which makes it trivially testable (section 13):</p>
<p>ts
// reduce.ts
export function reduce(
  cur: MatchState | undefined,
  next: MatchState
): MatchState {
  return cur &amp;&amp; cur.version &gt;= next.version ? cur : next;
}</p>
<p>And here is the same rule enforced atomically in Redis, so that two workers racing on the same match cannot corrupt each other. It also publishes the update so every gateway learns about it:</p>
<p>ts
// store.ts
import Redis from "ioredis";
import type { MatchState } from "./types";</p>
<p>export const redis = new Redis(process.env.REDIS_URL!);</p>
<p>const APPLY_IF_NEWER = <code>  local cur = redis.call('HGET', KEYS[1], 'v')   if cur and tonumber(cur) &gt;= tonumber(ARGV[1]) then     return 0   end   redis.call('HSET', KEYS[1], 'v', ARGV[1], 'json', ARGV[2])   redis.call('PUBLISH', 'match-updates', ARGV[2])   return 1</code>;</p>
<p>export async function applyIfNewer(
  s: MatchState,
  source: "ws" | "webhook" | "rest"
): Promise {
  const applied = await redis.eval(
    APPLY_IF_NEWER, 1,
    <code>match:${s.matchId}</code>, String(s.version), JSON.stringify(s)
  );
  metrics.count(applied ? "applied" : "rejected", { source, sport: s.sport });
  return applied === 1;
}</p>
<p>Two details worth noticing:</p>
<p>The source tag costs nothing and later answers a very valuable question: which channel delivered each update first? (Section 12.)
A Lua script runs atomically in Redis, so the read-compare-write can't interleave with another worker's. In production, register it with defineCommand so it's sent once and invoked by hash.</p>
<p>Caveat: a "highest version wins" reducer assumes each message carries full state (or cumulative state). If your provider streams deltas (for example, individual events), model events by ID with revisions and derive the score from them instead. That approach is covered in the previous articles in this series, and it still funnels through one idempotent writer.</p>
<ol>
<li>Channel 1: WebSocket ingest</li>
</ol>
<p>The socket is your primary live feed. Its job is simple: parse, normalize, apply. It must never do slow work in the message handler.</p>
<p>ts
// ingest-ws.ts
import WebSocket from "ws";
import { applyIfNewer } from "./store";
import { normalize } from "./normalize";   // provider payload -&gt; MatchState
import { reconcileAll } from "./reconcile";</p>
<p>const WS_URL = process.env.ORBISTATS_WS_URL!;   // placeholder: copy from the WebSocket docs
const KEY = process.env.ORBISTATS_KEY!;</p>
<p>let attempt = 0;</p>
<p>export function connect() {
  const ws = new WebSocket(WS_URL, { headers: { Authorization: <code>Bearer ${KEY}</code> } });
  let lastMsgAt = Date.now();</p>
<p>  const watchdog = setInterval(() =&gt; {
    if (Date.now() - lastMsgAt &gt; 30_000) ws.terminate();   // half-open socket guard
  }, 5_000);</p>
<p>  ws.on("open", async () =&gt; {
    attempt = 0;
    // subscribe payload is a placeholder: see the WebSocket API page
    ws.send(JSON.stringify({ action: "subscribe", sport: "football" }));
    await reconcileAll("reconnect");                       // heal any gap immediately
  });</p>
<p>  ws.on("message", (raw) =&gt; {
    lastMsgAt = Date.now();
    const state = normalize(JSON.parse(raw.toString()));
    if (state) void applyIfNewer(state, "ws");             // fire, don't block the loop
  });</p>
<p>  ws.on("close", () =&gt; {
    clearInterval(watchdog);
    const delay = Math.min(15_000, 500 * 2 ** attempt++) * (0.5 + Math.random() / 2);
    setTimeout(connect, delay);                            // exponential backoff + jitter
  });</p>
<p>  ws.on("error", (e) =&gt; console.error("ws error", e.message));
}</p>
<p>Three habits are baked in: a silence watchdog (TCP can look alive while delivering nothing), backoff with jitter (so a provider blip doesn't become a reconnect stampede), and a reconcile on every (re)connect so the gap is closed before you trust the stream again.</p>
<p>A bonus for resilience: because the reducer is idempotent, you can run two ingest instances in different regions, both applying the same stream. Duplicates are harmless, and failover time drops to zero. The cost is a second connection against your plan's limits, so check the pricing page first.</p>
<ol>
<li>Channel 2: Webhooks through a durable stream</li>
</ol>
<p>Webhooks are valuable precisely because they don't depend on a connection you hold open. But the receiving endpoint has one job: acknowledge fast, and never lose the event. So the HTTP handler does almost nothing, it verifies and writes to a durable Redis Stream, and a separate worker does the processing.</p>
<p>ts
// webhook-receiver.ts
import express from "express";
import crypto from "node:crypto";
import { redis } from "./store";</p>
<p>const app = express();
app.use("/webhooks/orbistats", express.raw({ type: "application/json" }));</p>
<p>app.post("/webhooks/orbistats", async (req, res) =&gt; {
  // header name + algorithm are placeholders: see the Webhooks API page
  const sig = req.header("x-signature") ?? "";
  const expected = crypto
    .createHmac("sha256", process.env.WEBHOOK_SECRET!)
    .update(req.body)
    .digest("hex");</p>
<p>  const ok =
    sig.length === expected.length &amp;&amp;
    crypto.timingSafeEqual(Buffer.from(sig), Buffer.from(expected));
  if (!ok) return res.sendStatus(401);</p>
<p>  // Durable enqueue FIRST, ack SECOND.
  await redis.xadd("webhook-events", "MAXLEN", "~", 100_000, "*", "payload", req.body.toString());
  res.sendStatus(200);
});</p>
<p>app.listen(3000);</p>
<p>Notice what's missing: there's no "have I seen this event ID?" check in the receiver. That's deliberate. A dedupe key written before the enqueue succeeds can swallow an event if the enqueue then fails and the provider retries. Dedupe is an optimization, the reducer is the guarantee, and it already makes duplicates harmless.</p>
<p>The worker reads from a consumer group, so work is shared across instances and unacknowledged messages can be reclaimed after a crash:</p>
<p>ts
// webhook-worker.ts
import { redis, applyIfNewer } from "./store";
import { normalize } from "./normalize";</p>
<p>const STREAM = "webhook-events", GROUP = "workers";
const CONSUMER = <code>worker-${process.pid}</code>;</p>
<p>await redis.xgroup("CREATE", STREAM, GROUP, "$", "MKSTREAM").catch(() =&gt; {}); // exists already</p>
<p>for (;;) {
  const res: any = await redis.xreadgroup(
    "GROUP", GROUP, CONSUMER, "COUNT", 50, "BLOCK", 5000, "STREAMS", STREAM, "&gt;"
  );
  if (!res) continue;</p>
<p>  for (const [, entries] of res) {
    for (const [id, fields] of entries) {
      try {
        const state = normalize(JSON.parse(fields[1]));
        if (state) await applyIfNewer(state, "webhook");
        await redis.xack(STREAM, GROUP, id);        // ack only after a successful apply
      } catch (err) {
        console.error("webhook processing failed", id, err);
        // left pending: a periodic XAUTOCLAIM job retries it, then dead-letters after N attempts
      }
    }
  }
}</p>
<p>Retry behaviour, payloads and signing rules belong to your provider, so read the Webhooks API page and the documentation before finalizing the receiver.</p>
<ol>
<li>Channel 3: REST for snapshots, reconciliation and cold data</li>
</ol>
<p>REST plays three roles in this design, and none of them is "poll for live scores".</p>
<p>Role 1: Snapshot on start and reconnect. Fill state before streaming resumes.</p>
<p>Role 2: Reconciliation. A periodic sweep compares the source of truth against your state:</p>
<p>ts
// reconcile.ts
import { applyIfNewer } from "./store";
import { normalize } from "./normalize";</p>
<p>const BASE = "<a href="https://api.orbistats.com/v1">https://api.orbistats.com/v1</a>";
const KEY = process.env.ORBISTATS_KEY!;</p>
<p>export async function reconcileAll(reason: "interval" | "reconnect" | "startup") {
  // One call for ALL live matches of a sport, if the endpoint returns a list (verify!).
  const res = await fetch(<code>${BASE}/football/matches/live</code>, {
    headers: { Authorization: <code>Bearer ${KEY}</code> },
  });
  if (res.status === 429) { metrics.count("rate_limited"); return; }
  if (!res.ok) { metrics.count("reconcile_error", { status: res.status }); return; }</p>
<p>  const body = await res.json();
  const items = Array.isArray(body) ? body : body.matches ?? [body];</p>
<p>  for (const raw of items) {
    const state = normalize(raw);
    if (!state) continue;
    const applied = await applyIfNewer(state, "rest");
    // If REST changed something, a push channel missed it. That's DRIFT.
    if (applied) metrics.count("drift_detected", { sport: state.sport, reason });
  }
}</p>
<p>setInterval(() =&gt; reconcileAll("interval").catch(console.error), 30_000);</p>
<p>That drift_detected counter is one of the most useful numbers in the whole system. If it's always zero, your push channels are healthy. If it spikes, you've found an incident before your users did.</p>
<p>Role 3: Cold data. Fixtures, standings, team metadata and history change slowly and cache beautifully. See the Sports Data API and the Historical Sports Data API.</p>
<p>The request-budget math</p>
<p>Plans have request limits, and reconciliation can quietly eat them. The formula is simple:</p>
<p>text
requests/day = (86,400 / interval_seconds) × calls_per_sweep
Strategy	Calls per sweep	Interval	Requests/day
One list call for all live matches	1	30 s	2,880
One call per live match, 10 matches	10	30 s	28,800
One call per live match, 10 matches	10	60 s	14,400</p>
<p>The takeaway: prefer list endpoints over per-match calls, because the cost scales with sweeps, not with match count. Against a free tier capped at 150 requests per day (per the FAQ; confirm current limits on the pricing page), even the cheapest row above is out of reach, so continuous reconciliation is a paid-tier design. On a free key, reconcile only at reconnect and startup.</p>
<ol>
<li>The cache hierarchy</li>
</ol>
<p>Caching in this pipeline has three tiers, and the rule for each tier is different:</p>
<p>text
L1  in-process Map        microseconds   push-updated, no TTL
L2  Redis                 sub-millisecond-ish, same region   source of truth for current state
L3  HTTP / CDN            for slow-changing data only         TTL + stale-while-revalidate
Match TTLs by data class</p>
<p>These are my starting points, not provider specifications:</p>
<p>Data class	Where cached	Policy
Live score / live odds	L1 + L2	Push-updated. No TTL. Never put a CDN in front without a deliberate micro-TTL
Pre-match odds	L2 / short CDN	Seconds to a minute, depending on how fast you need them
Fixtures	CDN + L2	Minutes (they change when schedules change)
Standings	CDN + L2	A minute or two during matches; longer between matchdays
Finished-match results	CDN + L2	Long, but allow purge on correction
Historical seasons	CDN	Very long. Closed seasons are effectively immutable
Stale-while-revalidate with request coalescing</p>
<p>For the slow-changing tiers, two tricks matter. Stale-while-revalidate serves slightly old data instantly while refreshing in the background. Request coalescing (often called singleflight) makes sure that 500 simultaneous cache misses cause one upstream request, not 500. Without it, a cache expiry at the wrong moment becomes a thundering herd that burns your rate limit.</p>
<p>ts
// swr-cache.ts
type Entry = { value: T; freshUntil: number; staleUntil: number };</p>
<p>const cache = new Map&lt;string, Entry&gt;();
const inflight = new Map&lt;string, Promise&gt;();</p>
<p>function refresh(key: string, ttlMs: number, staleMs: number, loader: () =&gt; Promise): Promise {
  let p = inflight.get(key) as Promise | undefined;
  if (!p) {
    p = loader()
      .then((value) =&gt; {
        const now = Date.now();
        cache.set(key, { value, freshUntil: now + ttlMs, staleUntil: now + ttlMs + staleMs });
        return value;
      })
      .finally(() =&gt; inflight.delete(key));
    inflight.set(key, p);        // everyone arriving during the fetch shares this promise
  }
  return p;
}</p>
<p>export async function cached(
  key: string, ttlMs: number, staleMs: number, loader: () =&gt; Promise
): Promise {
  const e = cache.get(key) as Entry | undefined;
  const now = Date.now();</p>
<p>  if (e &amp;&amp; now &lt; e.freshUntil) return e.value;                       // fresh hit
  if (e &amp;&amp; now &lt; e.staleUntil) {
    refresh(key, ttlMs, staleMs, loader).catch(() =&gt; {});            // serve stale, refresh behind the scenes
    return e.value;
  }
  return refresh(key, ttlMs, staleMs, loader);                       // miss: one shared fetch
}</p>
<p>// usage
// const standings = await cached("standings:epl", 60_000, 300_000, () =&gt; fetchStandings("epl"));</p>
<p>For your own HTTP API, express the same idea in headers so CDNs and browsers cooperate:</p>
<p>ts
// slow data: let the CDN absorb the load
res.set("Cache-Control", "public, s-maxage=60, stale-while-revalidate=300");</p>
<p>// live data: never let an intermediary cache it by accident
res.set("Cache-Control", "no-store");</p>
<p>A forgotten max-age on a live endpoint looks exactly like provider latency. If you've ever chased phantom lag, this is often the culprit.</p>
<ol>
<li>The gateway: L1 cache, snapshot-on-connect, fan-out</li>
</ol>
<p>Gateways are the processes that hold client WebSocket connections. Each one keeps an L1 cache populated by Redis pub/sub, so serving a new client never touches Redis on the hot path.</p>
<p>Redis pub/sub is fire-and-forget: if a gateway's subscription drops, it misses messages. So the gateway warms from Redis after every (re)subscribe, and compares versions so a stale warm-up can't overwrite a newer pushed update.</p>
<p>ts
// gateway.ts
import Redis from "ioredis";
import { WebSocketServer, WebSocket } from "ws";
import { redis } from "./store";</p>
<p>type Cached = { v: number; json: string };
const l1 = new Map&lt;string, Cached&gt;();                    // matchId -&gt; latest
const subs = new Map&lt;string, Set&gt;();          // matchId -&gt; subscribers</p>
<p>function put(matchId: string, v: number, json: string): boolean {
  const cur = l1.get(matchId);
  if (cur &amp;&amp; cur.v &gt;= v) return false;                   // same version rule as the reducer
  l1.set(matchId, { v, json });
  return true;
}</p>
<p>function fanout(matchId: string, json: string) {
  for (const ws of subs.get(matchId) ?? []) {
    if (ws.readyState === WebSocket.OPEN) ws.send(json);
  }
}</p>
<p>async function warmFromRedis() {
  for await (const keys of redis.scanStream({ match: "match:*", count: 200 })) {
    const pipe = redis.pipeline();
    for (const k of keys as string[]) pipe.hmget(k, "v", "json");
    const rows = (await pipe.exec()) ?? [];
    for (const [, row] of rows) {
      const [v, json] = row as [string | null, string | null];
      if (v &amp;&amp; json) put(JSON.parse(json).matchId, Number(v), json);
    }
  }
}</p>
<p>const sub = new Redis(process.env.REDIS_URL!);
sub.on("message", (_ch, json) =&gt; {
  const s = JSON.parse(json);
  if (put(s.matchId, s.version, json)) fanout(s.matchId, json);
});
sub.on("ready", async () =&gt; {                            // fires on every reconnect too
  await sub.subscribe("match-updates");
  await warmFromRedis();                                 // heal the pub/sub gap
});</p>
<p>const wss = new WebSocketServer({ port: 8080 });
wss.on("connection", (ws) =&gt; {
  ws.on("message", (raw) =&gt; {
    const { subscribe = [] } = JSON.parse(raw.toString());
    for (const id of subscribe as string[]) {
      if (!subs.has(id)) subs.set(id, new Set());
      subs.get(id)!.add(ws);
      const snap = l1.get(id);
      if (snap) ws.send(snap.json);                      // snapshot first, deltas after
    }
  });
  ws.on("close", () =&gt; {
    for (const set of subs.values()) set.delete(ws);
  });
});</p>
<p>For slow clients, conflate rather than queue: only the latest state matters for a score, so replace pending messages for the same match instead of buffering all of them. The fan-out and backpressure details are in the previous article, The Hidden Latency Stack style of analysis, and the principle is the same here.</p>
<p>If you don't want to own the client side at all, the drop-in widgets cover rendering for common cases.</p>
<ol>
<li>Failure modes and the degradation ladder</li>
</ol>
<p>Write this table before an incident, not during one:</p>
<p>What fails	What still works	What users see	Recovery
WebSocket drops	Webhooks + REST reconcile	Slightly delayed updates (up to the reconcile interval)	Backoff reconnect, then snapshot heal
Webhook endpoint down	WebSocket + REST	Nothing visible. Side effects may lag	Provider retries, worker drains the stream
REST rate-limited (429)	WebSocket + webhooks	Nothing visible	Back off, rely on push, reduce sweep frequency
Redis unavailable	Ingest buffers or retries	Last-known values from gateway L1	Restore Redis, warm L1, replay
One gateway crashes	Other gateways	Clients reconnect and get a snapshot	Load balancer drains, client reconnect
Provider incident	Your cache	Frozen, but labeled data	Show a "delayed" indicator</p>
<p>That last row deserves a design decision: make staleness visible. If no update has arrived for a live match within some window, mark it ("updating…") rather than silently showing an old score as current. Users forgive a delay they can see far more than a wrong score they can't. For provider-side incidents, point your on-call at the status page so you can quickly tell "their side" from "ours".</p>
<ol>
<li>Tuning per sport</li>
</ol>
<p>Thirteen sports means thirteen different traffic shapes. One set of constants won't fit all of them. These are my starting heuristics, which you should tune with your own data:</p>
<p>Sport type	Update character	Reconcile interval (starting point)
Football	Sparse events, long quiet stretches	30–60 s
Basketball	Dense scoring, constant change	15–30 s
Tennis	Point-by-point	15–30 s
Cricket	Ball-by-ball, very long matches	30–60 s
Golf	Many simultaneous competitors, slow ticks	60 s+
Horse racing	Short bursts, long gaps	Event-driven plus a slow sweep</p>
<p>Make these config, not code:</p>
<p>ts
// profiles.ts
export const PROFILES: Record&lt;string, { reconcileMs: number; silenceMs: number }&gt; = {
  football:   { reconcileMs: 30_000, silenceMs: 45_000 },
  basketball: { reconcileMs: 15_000, silenceMs: 20_000 },
  tennis:     { reconcileMs: 15_000, silenceMs: 30_000 },
  golf:       { reconcileMs: 60_000, silenceMs: 120_000 },
};</p>
<p>The silenceMs value drives your "no update for too long" alert. A 45-second silence in basketball is alarming, but in golf it's routine.</p>
<ol>
<li>Observability: the metrics that matter</li>
</ol>
<p>You can run this pipeline with a handful of metrics. Each answers a specific question:</p>
<p>Metric	Question it answers
applied{source} / rejected{source}	Which channel is actually doing the work? Is anything redundant?
drift_detected{sport,reason}	Are push channels missing updates?
staleness_seconds{sport} (age of freshest data per live match)	Is anything silently stuck?
Stream lag and oldest-pending age	Is the webhook worker keeping up?
ingest_to_emit_ms p50/p95/p99	Is our own hot path healthy?
rate_limited	Are we spending request budget faster than planned?</p>
<p>The most interesting one is the channel race: for each update, record which source applied it first. Over a week it tells you, with your own data, how much each channel contributes. If the WebSocket wins nearly every time, webhooks are your insurance, not your speed path. That's useful evidence for deciding what you actually need on your plan.</p>
<p>Here's the staleness sampler:</p>
<p>ts
// staleness.ts
const lastApplied = new Map&lt;string, number&gt;();    // matchId -&gt; Date.now() of last apply
export function noteApplied(matchId: string) { lastApplied.set(matchId, Date.now()); }</p>
<p>setInterval(() =&gt; {
  const now = Date.now();
  for (const [id, t] of lastApplied) {
    const profile = PROFILES[sportOf(id)];
    const age = now - t;
    metrics.gauge("staleness_seconds", age / 1000, { sport: sportOf(id) });
    if (isLive(id) &amp;&amp; age &gt; profile.silenceMs) alert(<code>match ${id} silent for ${Math.round(age / 1000)}s</code>);
  }
}, 5_000);</p>
<p>Always report percentiles, never averages. The reasoning (why tails dominate user experience and why you can't average percentiles) is in the previous articles.</p>
<ol>
<li>Chaos testing the reducer</li>
</ol>
<p>Everything above rests on one claim: the reducer gives the same final state no matter how messages arrive. Don't trust it, prove it. Because reduce is pure, a randomized test is cheap:</p>
<p>ts
// reduce.test.ts
import assert from "node:assert/strict";
import { reduce } from "./reduce";
import type { MatchState } from "./types";</p>
<p>const mk = (version: number): MatchState =&gt; ({
  matchId: "m1", sport: "football", version, status: "live", minute: version,
  home: { name: "A", score: Math.floor(version / 10) },
  away: { name: "B", score: Math.floor(version / 20) },
});</p>
<p>const shuffle = &lt;T,&gt;(a: T[]): T[] =&gt; {
  const r = [...a];
  for (let i = r.length - 1; i &gt; 0; i--) {
    const j = Math.floor(Math.random() * (i + 1));
    [r[i], r[j]] = [r[j], r[i]];
  }
  return r;
};</p>
<p>const stream = Array.from({ length: 50 }, (_, i) =&gt; mk(i + 1));
const expected = stream[stream.length - 1];</p>
<p>for (let run = 0; run &lt; 1000; run++) {
  // duplicates (30% of events delivered twice) + arbitrary reordering
  const chaos = shuffle([...stream, ...stream.filter(() =&gt; Math.random() &lt; 0.3)]);</p>
<p>  let state: MatchState | undefined;
  for (const ev of chaos) state = reduce(state, ev);</p>
<p>  assert.deepEqual(state, expected);
}
console.log("1000 chaos runs: converged to the same final state every time");</p>
<p>Then extend it to drops: delete a random 20% of events, apply them, then run one "REST snapshot" containing the final state, and assert convergence. That test is exactly the guarantee you're buying with reconciliation: drops are healed, duplicates are harmless, reordering is irrelevant.</p>
<p>For integration-level confidence, record a real session of the WebSocket feed to a file, then replay it with injected delays, duplicates and disconnects against a staging copy of the pipeline. Practicing against real payloads starts with a free key, so see the Quickstart and try routes in the Sandbox first.</p>
<ol>
<li>Capacity, cost and limits</li>
</ol>
<p>A few practical budgeting points:</p>
<p>Plans: The pricing page lists Free ($0), Starter ($19/mo, positioned for real-time production), Growth ($79/mo, positioned around historical data, widgets and webhooks) and Enterprise (custom, dedicated infrastructure). Confirm which tier includes which delivery methods before you design around them, on the pricing page.
Connection counts: one upstream socket per ingest instance, not one per user. That's the main reason to put your own gateway tier in front of clients.
Memory: an L1 entry per live match is tiny. Even thousands of concurrent matches fit comfortably in a gateway's heap, which is why a push-updated in-process cache is so cheap.
Redis: pub/sub fan-out to a handful of gateways is light. The thing to watch is the webhook stream length, so cap it with MAXLEN as shown.
Fan-out cost: it scales with matches × subscribers × update rate. Conflation and diff-before-send are your levers.
15. Checklist and next steps
 Every channel writes through one atomic applyIfNewer
 Reducer is pure, idempotent and covered by a chaos test
 WebSocket: silence watchdog, jittered backoff, reconcile on reconnect
 Webhooks: verify signature, durable enqueue, then ack, worker with consumer group and dead-lettering
 REST: list endpoints over per-match calls, 429 handling, drift counter
 Live data is push-updated; TTLs only for slow data
 Singleflight on every REST-backed cache
 Gateways warm L1 after every pub/sub reconnect
 Per-sport profiles in config
 Staleness is visible to users and alerted on
 Degradation ladder written down before the first incident</p>
<p>The core idea fits in one sentence: let every channel be unreliable, and make the one place that merges them impossible to confuse.</p>
<p>Build it yourself:</p>
<p>Create a free key on the signup page.
Read the documentation and the API reference to map real payload fields onto the MatchState type.
Explore the delivery options: WebSocket API, Webhooks API, Live Scores API.
If you're building for price-sensitive users, start with the Odds API and the Trading Desks use case.
Browse everything else at Orbistats.</p>
<p>What does your reconciliation interval look like today, and have you ever measured how often your push channel actually drops updates? Share your drift numbers in the comments. I'd love to compare.</p>
]]></content:encoded></item><item><title><![CDATA[The Hidden Latency Stack: What Happens Between a Goal and Your Screen]]></title><description><![CDATA[A striker hits the ball. Forty thousand people in the stadium scream. Somewhere, a phone buzzes with a score update.
How long did that take, and where did the time go?
Most developers answer with one ]]></description><link>https://bettechmagnetics.hashnode.dev/the-hidden-latency-stack-what-happens-between-a-goal-and-your-screen</link><guid isPermaLink="true">https://bettechmagnetics.hashnode.dev/the-hidden-latency-stack-what-happens-between-a-goal-and-your-screen</guid><category><![CDATA[Web Development]]></category><category><![CDATA[APIs]]></category><category><![CDATA[distributed systems]]></category><category><![CDATA[websockets]]></category><category><![CDATA[performance]]></category><dc:creator><![CDATA[Vijay Choudhary]]></dc:creator><pubDate>Tue, 06 Oct 2026 16:31:20 GMT</pubDate><enclosure url="https://cdn.hashnode.com/uploads/covers/6a47ac8dec0b8a17b4dc915d/248abaf2-2125-4ec9-9258-bd087ec680ff.png" length="0" type="image/jpeg"/><content:encoded><![CDATA[<p>A striker hits the ball. Forty thousand people in the stadium scream. Somewhere, a phone buzzes with a score update.</p>
<p>How long did that take, and where did the time go?</p>
<p>Most developers answer with one number: "the API latency". But a live score isn't delivered by one system. It crosses at least eight layers, owned by different parties, each with its own physics, queues and failure modes. This article walks the whole stack top to bottom, shows where time hides at each layer, and gives you code to measure the parts you control.</p>
<p>I'll use the Orbistats sports-data API as the concrete example for the parts you integrate with, since it exposes REST, WebSocket and Webhooks behind one normalized schema across 13 sports. Where I describe what happens inside a provider, I'm describing a typical industry architecture, not Orbistats' internals.</p>
<p>No fake benchmarks. You won't find invented millisecond results here. Where I quote a number, it's either a physical constant, a mathematical result, or something you can compute yourself with the code provided.</p>
<p>Table of contents
The stack at a glance
Layer 0: The physical floor (you can't beat light)
Layer 1: Capturing the event
Layer 2: Validation, correction and the "provisional goal" problem
Layer 3: Ingestion, normalization and ordering
Layer 4: Fan-out and the queueing math nobody mentions
Layer 5: Transport (TCP, TLS, WebSocket upgrade)
Layer 6: Your backend
Layer 7: The client (event loop, frames and radios)
Tail latency: why p99 is the real product
Putting it together: a tracing harness
What to ask your data provider
Reference architecture and next steps</p>
<ol>
<li>The stack at a glance
text
 ┌─────────────────────────────────────────────────────────┐
 │ L0  Physics         distance / speed of light in fiber  │
 ├─────────────────────────────────────────────────────────┤
 │ L1  Capture         the event is observed and entered   │
 ├─────────────────────────────────────────────────────────┤
 │ L2  Validation      confirmed? corrected? (VAR, review) │
 ├─────────────────────────────────────────────────────────┤
 │ L3  Ingestion       parse, normalize, dedupe, order     │
 ├─────────────────────────────────────────────────────────┤
 │ L4  Fan-out         publish to N subscribers            │
 ├─────────────────────────────────────────────────────────┤
 │ L5  Transport       TCP/TLS/WebSocket/HTTP              │
 ├─────────────────────────────────────────────────────────┤
 │ L6  Your backend    cache, enrich, re-broadcast         │
 ├─────────────────────────────────────────────────────────┤
 │ L7  Client          parse, state, layout, paint, display│
 └─────────────────────────────────────────────────────────┘</li>
</ol>
<p>Providers usually quote latency for L3 to L5 only. Orbistats, for example, positions its live feed as sub-50ms and discusses what sub-50ms actually requires, end to end. That's a statement about their segment. L0, L1, L2 and L6 to L7 sit outside any provider's control, and they are frequently larger.</p>
<ol>
<li>Layer 0: The physical floor (you can't beat light)</li>
</ol>
<p>Light in optical fiber travels at roughly two-thirds of its vacuum speed, about 200,000 km/s, which is 200 km per millisecond. Signals also don't travel in straight lines, so real paths are longer than the great-circle distance. That makes the straight-line figure a hard lower bound.</p>
<p>For London to Singapore (about 10,900 km apart):</p>
<p>text
one-way floor ≈ 10,900 km / 200 km/ms ≈ 54 ms
round-trip floor ≈ 108 ms</p>
<p>That round trip alone is already over 100 ms, before a single router, queue or line of code is involved. Which tells you something useful: "sub-100ms" is only meaningful relative to where the data is served from. Regional edges and servers close to your users aren't an optimization, they're a prerequisite.</p>
<p>Here's a tiny calculator to find the floor for your own user geography:</p>
<p>ts
// floor.ts
const FIBER_KM_PER_MS = 200; // ~2/3 c</p>
<p>type Pt = { lat: number; lon: number };</p>
<p>function haversineKm(a: Pt, b: Pt): number {
  const R = 6371;
  const toRad = (d: number) =&gt; (d * Math.PI) / 180;
  const dLat = toRad(b.lat - a.lat);
  const dLon = toRad(b.lon - a.lon);
  const h =
    Math.sin(dLat / 2) ** 2 +
    Math.cos(toRad(a.lat)) * Math.cos(toRad(b.lat)) * Math.sin(dLon / 2) ** 2;
  return 2 * R * Math.asin(Math.sqrt(h));
}</p>
<p>export function floorRttMs(a: Pt, b: Pt): number {
  return (2 * haversineKm(a, b)) / FIBER_KM_PER_MS;
}</p>
<p>// Example: where is your server vs where are your users?
const london: Pt = { lat: 51.5, lon: -0.12 };
const singapore: Pt = { lat: 1.35, lon: 103.82 };
console.log(<code>RTT floor: ${floorRttMs(london, singapore).toFixed(0)} ms</code>);</p>
<p>Compare that floor with what you actually measure (section 11). The gap between them is the part software can improve.</p>
<ol>
<li>Layer 1: Capturing the event</li>
</ol>
<p>Nothing in a sports-data system starts with an API. It starts with a human or a sensor observing something. Depending on the sport and provider, the source may be on-site data collectors, official league data feeds, optical or wearable tracking, or a combination.</p>
<p>The important engineering consequence: capture latency is not zero and not constant. It varies with the sport, the competition tier and the source. A top-tier match with an official feed can behave very differently from a lower-division fixture covered by a single operator. If you serve many sports, expect a different baseline for each. A provider covering 13 sports, as Orbistats lists across football and its other sport pages, has to deal with 13 different capture realities.</p>
<p>What this means for you: don't expect one universal "goal-to-screen" number. Ask which sports and competitions have the fastest sources, and measure per sport.</p>
<ol>
<li>Layer 2: Validation, correction and the "provisional goal" problem</li>
</ol>
<p>Here's a problem pure speed obsession ignores: the first report isn't always right.</p>
<p>A goal is entered, then disallowed after a video review.
A penalty is awarded, then overturned.
A tennis point is credited to the wrong player and fixed seconds later.
A cricket run is re-attributed (say, a leg-bye reclassified).</p>
<p>Any live-score system lives on a three-way tradeoff:</p>
<p>text
        Speed
         /<br />        /  <br />       /    <br />      /______<br /> Correctness  Stability</p>
<p>Publish instantly and you'll sometimes show a goal that gets removed. Wait for confirmation and you're slower. Neither is wrong; it's a product decision. But your client code must handle corrections gracefully, or users see a score that goes 1–0, 2–0, then back to 1–0 with no explanation.</p>
<p>The robust approach is to model events, not just the score, and derive the score from the events:</p>
<p>ts
// state.ts - illustrative model; real field names come from your provider's docs
type Status = "provisional" | "confirmed" | "cancelled";</p>
<p>interface MatchEvent {
  id: string;          // stable per real-world event
  matchId: string;
  seq: number;         // increases each time THIS event is revised
  type: "goal" | "card" | "sub";
  team?: "home" | "away";
  status: Status;
}</p>
<p>class MatchState {
  private events = new Map&lt;string, MatchEvent&gt;();</p>
<p>  // Idempotent + order-safe: duplicates and late arrivals can't corrupt state
  apply(ev: MatchEvent): boolean {
    const cur = this.events.get(ev.id);
    if (cur &amp;&amp; cur.seq &gt;= ev.seq) return false; // stale or duplicate
    this.events.set(ev.id, ev);
    return true;
  }</p>
<p>  score() {
    let home = 0, away = 0;
    for (const e of this.events.values()) {
      if (e.type !== "goal" || e.status === "cancelled") continue;
      if (e.team === "home") home++;
      else if (e.team === "away") away++;
    }
    return { home, away };
  }</p>
<p>  hasProvisional() {
    return [...this.events.values()].some(e =&gt; e.status === "provisional");
  }
}</p>
<p>Two payoffs. First, applying the same update twice (very common with retries) is harmless. Second, a cancellation just changes one event's status and the score recomputes correctly. In the UI you can show a subtle "pending review" badge while hasProvisional() is true, which turns a correctness problem into a feature.</p>
<p>Check your provider's API reference for how events, revisions and statuses are actually represented before copying this model.</p>
<ol>
<li>Layer 3: Ingestion, normalization and ordering</li>
</ol>
<p>Once an event exists, the provider has to turn many differently-shaped inputs into one clean schema. This is the same "normalization" idea that makes the Orbistats Odds API useful: you consume one consistent shape instead of maintaining a parser per source.</p>
<p>Typical work at this layer:</p>
<p>Parse the incoming feed message.
Normalize team, player and competition IDs and field names.
Deduplicate, since sources frequently resend.
Order, since sources deliver out of order under load.
Persist and publish.</p>
<p>Each step is cheap in isolation. The danger is ordering. If two updates for the same match are processed on different workers, they can be published in the wrong order. The standard fix is partitioning by key: route all messages for one match to the same worker/partition so per-match order is preserved, while different matches proceed in parallel.</p>
<p>text
   incoming feed
        │
   hash(match_id) % N
   ┌────┼────┬────┐
   ▼    ▼    ▼    ▼
  W0   W1   W2   W3     ← one match always lands on the same worker
   └────┴────┴────┘
        │
     publish (ordered per match)</p>
<p>If you build your own ingest (for example, receiving webhooks and re-broadcasting), apply the same rule on your side.</p>
<ol>
<li>Layer 4: Fan-out and the queueing math nobody mentions</li>
</ol>
<p>One goal must reach many consumers: your app, thousands of other customers, dashboards, caches. That's fan-out, and it's where queues enter the picture.</p>
<p>Queueing theory has a result every engineer who cares about latency should memorize. For a simple single-server queue (M/M/1), the average time a message spends in the system is:</p>
<p>text
W = (1/μ) / (1 − ρ)</p>
<p>where 1/μ is the service time and ρ is utilization (the fraction of capacity in use).</p>
<p>Utilization ρ	Time in system vs. service time
50%	2×
80%	5×
90%	10×
95%	20×
99%	100×</p>
<p>Latency doesn't grow linearly with load, it explodes near saturation. This explains an infuriating real-world pattern: everything is fast all week, then on a packed Saturday with dozens of simultaneous matches and goals, the same code on the same hardware becomes slow. Nothing broke. Utilization crossed a knee.</p>
<p>Practical consequences:</p>
<p>Keep headroom. A system running at 70% on a quiet day may be at 95% during peak.
Scale on queue age (how old is the oldest waiting message), not just CPU.
Set explicit limits and decide what to do when full: drop, conflate or reject. The wrong default is "buffer forever", which turns overload into ever-growing latency.
Slow consumers and conflation</p>
<p>In a fan-out gateway, one slow client can drag down memory and, in a naive design, everyone else. Live scores have a lovely property: only the latest state matters. So when a client can't keep up, you don't queue every intermediate update, you conflate them into the newest one.</p>
<p>js
// gateway.js (Node, using the 'ws' library)
const MAX_BUFFERED = 64 * 1024; // bytes waiting in the socket buffer</p>
<p>class ClientConn {
  constructor(ws) {
    this.ws = ws;
    this.pending = new Map();   // key -&gt; latest message (conflated)
    this.flushing = false;
  }</p>
<p>  send(key, msg) {
    this.pending.set(key, msg);          // newer replaces older for the same key
    if (!this.flushing) this.scheduleFlush();
  }</p>
<p>  scheduleFlush() {
    this.flushing = true;
    setImmediate(() =&gt; {
      this.flushing = false;
      // Slow consumer: don't pile more bytes onto a clogged socket.
      if (this.ws.bufferedAmount &gt; MAX_BUFFERED) {
        this.scheduleFlushLater();
        return;
      }
      for (const msg of this.pending.values()) {
        this.ws.send(JSON.stringify(msg));
      }
      this.pending.clear();
    });
  }</p>
<p>  scheduleFlushLater() {
    setTimeout(() =&gt; this.scheduleFlush(), 50);
  }
}</p>
<p>The slow client gets fewer, fresher updates instead of a growing backlog of stale ones. That is always the right trade for scores. (It's the wrong trade for something like a ledger, where every message matters. Know your data.)</p>
<ol>
<li>Layer 5: Transport (TCP, TLS, WebSocket upgrade)</li>
</ol>
<p>Every protocol choice has a round-trip cost.</p>
<p>A cold HTTPS request typically needs:</p>
<p>text
1 RTT  TCP handshake
1 RTT  TLS 1.3 handshake  (resumption can reduce this)
1 RTT  the HTTP request/response itself
────────
≈ 3 RTT before you get your first byte</p>
<p>A cold WebSocket needs all of that plus the HTTP upgrade round trip. But you pay it once, after which every update is a one-way push with no per-message handshake. This is the real reason streaming beats polling for live data: it's not that WebSocket frames are magically faster, it's that polling pays handshake and request overhead repeatedly and adds an interval-sized staleness window. You can see the options side by side on the WebSocket API and Live Scores API pages.</p>
<p>A few transport details that bite in practice:</p>
<p>Connection reuse. A fresh connection per poll multiplies the handshake cost above. Use keep-alive.
Nagle's algorithm. TCP may delay small packets to batch them. For small, latency-sensitive messages, set TCP_NODELAY (socket.setNoDelay(true) in Node). Most WebSocket libraries already do this, but verify.
Head-of-line blocking. TCP delivers bytes in order. A single lost packet stalls everything behind it until retransmission. On lossy mobile networks, this causes the "frozen then catches up" feel.
Compression trade-off. permessage-deflate shrinks payloads but costs CPU and adds latency per message. For tiny score updates, it's often a net loss. Measure before enabling.
Payload size. Smaller messages serialize, transmit and parse faster. Measure what you actually send:
js
const verbose = { match_id: "match_50231", status: "live", minute: 72,
  home: { name: "Manchester City", score: 2 }, away: { name: "Arsenal", score: 1 } };
const compact = ["match_50231", "live", 72, 2, 1];</p>
<p>console.log("verbose bytes:", Buffer.byteLength(JSON.stringify(verbose)));
console.log("compact bytes:", Buffer.byteLength(JSON.stringify(compact)));</p>
<p>A common production pattern: send the verbose "entity" data once (team names, IDs), then stream only compact deltas.</p>
<ol>
<li>Layer 6: Your backend</li>
</ol>
<p>If you re-broadcast provider data to your own users (and you should, rather than exposing your key in browsers), you add a layer. Everything in layers 3 to 5 applies again, on your hardware:</p>
<p>text
Provider WS ──► ingest ──► state store ──► gateway ──► browsers
                   │            │
              (parse, diff)   (Redis / memory)</p>
<p>Rules of thumb:</p>
<p>One upstream connection, many downstream. Don't open a provider socket per user.
Diff before broadcasting. If the payload didn't change, don't send it.
Keep the hot path tiny. No database round trips between "message received" and "message emitted". Write to storage asynchronously.
Snapshot plus stream. New clients get the current state from the store first, then live deltas, so they never wait for the next change to see anything.
Resync after reconnect. After any upstream drop, pull a REST snapshot to heal gaps. Endpoints and parameters are in the documentation.
9. Layer 7: The client (event loop, frames and radios)</p>
<p>The bytes arrived. You're not done.</p>
<p>The JavaScript event loop. Your message handler runs on the same thread as rendering. A 40 ms task blocks the next frame. Keep handlers tiny, and push heavy work elsewhere.</p>
<p>Frame budget. At 60 Hz a frame is about 16.7 ms. A state change right after a frame starts may not appear until the next one, so display adds up to one frame of delay. Batch updates into requestAnimationFrame so you render once per frame, not once per message:</p>
<p>js
const latest = new Map();
let raf = 0;</p>
<p>export function enqueue(tick) {
  latest.set(tick.match_id, tick);
  if (!raf) {
    raf = requestAnimationFrame(() =&gt; {
      raf = 0;
      for (const t of latest.values()) paint(t);
      latest.clear();
    });
  }
}</p>
<p>Background tabs. Browsers throttle timers and pause requestAnimationFrame in hidden tabs. Resync on return:</p>
<p>js
document.addEventListener("visibilitychange", () =&gt; {
  if (document.visibilityState === "visible") resyncFromSnapshot();
});</p>
<p>Mobile radios. Cellular modems idle to save power, and waking one adds delay to the first packet after quiet periods. A persistent connection with keepalives can keep the path warm, at some battery cost. This is another reason a stream often feels snappier than polling on phones.</p>
<p>Garbage collection and layout. Allocating a new object per tick and re-rendering a long list will eventually cause pauses. Update only the DOM nodes that changed, and virtualize long lists.</p>
<p>If you'd rather not own this layer at all, the drop-in widgets render live scores and match centers for you.</p>
<ol>
<li>Tail latency: why p99 is the real product</li>
</ol>
<p>Averages lie. What users remember is the worst moments. Two facts make tails matter more than they seem.</p>
<p>Fact 1: tails compound across calls. If one request has a 1% chance of being slow, a page that makes N independent calls hits at least one slow response with probability 1 − 0.99^N:</p>
<p>js
const pSlow = (p, n) =&gt; 1 - Math.pow(1 - p, n);
console.log(pSlow(0.01, 1));    // 1%
console.log(pSlow(0.01, 10));   // ~9.6%
console.log(pSlow(0.01, 100));  // ~63.4%</p>
<p>A dashboard with 100 live widgets is usually experiencing someone's p99.</p>
<p>Fact 2: you can't average percentiles. The mean of p99 values across servers isn't the fleet p99. Record raw samples in a histogram and compute percentiles from that.</p>
<p>Node has a built-in histogram, so you don't need a dependency:</p>
<p>js
import { createHistogram } from "node:perf_hooks";</p>
<p>const hops = {
  ingest_to_emit: createHistogram(),
  emit_to_recv:   createHistogram(),
  recv_to_paint:  createHistogram(),
};</p>
<p>export function record(name, ms) {
  // Histograms need integers &gt;= 1; use microseconds for resolution
  hops[name].record(Math.max(1, Math.round(ms * 1000)));
}</p>
<p>export function snapshot() {
  const out = {};
  for (const [name, h] of Object.entries(hops)) {
    out[name] = {
      p50_ms: h.percentile(50) / 1000,
      p95_ms: h.percentile(95) / 1000,
      p99_ms: h.percentile(99) / 1000,
      max_ms: h.max / 1000,
      count:  h.count,
    };
  }
  return out;
}
11. Putting it together: a tracing harness</p>
<p>Now combine everything into one trace that follows a single event through your stack. Use monotonic time (performance.now()) for durations within one process, because wall-clock time can jump when NTP adjusts it. Use wall-clock time only to compare across machines, and sync those machines with NTP/chrony.</p>
<p>js
// trace.js
export function startTrace(msg) {
  msg._trace = {
    id: crypto.randomUUID(),
    ingest_wall: Date.now(),          // cross-machine reference
    ingest_mono: performance.now(),   // in-process duration reference
  };
  return msg;
}</p>
<p>export function markEmit(msg) {
  const t = msg._trace;
  t.emit_wall = Date.now();
  t.ingest_to_emit_ms = performance.now() - t.ingest_mono; // clean, same process
  return msg;
}
js
// browser
socket.onmessage = (e) =&gt; {
  const msg = JSON.parse(e.data);
  const recvWall = Date.now() + clockOffset;       // corrected using an offset handshake
  const emitToRecv = recvWall - msg._trace.emit_wall;</p>
<p>  requestAnimationFrame(() =&gt; {
    paint(msg);
    const recvToPaint = performance.now() - recvMono;
    navigator.sendBeacon("/metrics", JSON.stringify({
      id: msg._trace.id,
      ingest_to_emit: msg._trace.ingest_to_emit_ms,
      emit_to_recv: emitToRecv,
      recv_to_paint: recvToPaint,
    }));
  });
};</p>
<p>For the browser clock offset, estimate it NTP-style against a /time endpoint on your server: sample several times and trust the sample with the smallest round trip. Without this step, emit_to_recv mixes real network delay with clock disagreement and can even come out negative.</p>
<p>Then build one table, per sport and per region, and compare it to the floor from section 2:</p>
<p>Segment	p50	p95	p99	Physical floor
Provider → your ingest	?	?	?	RTT/2 to provider
Ingest → emit	?	?	?	~0
Emit → client	?	?	?	RTT/2 to user
Recv → paint	?	?	?	≤ 1 frame</p>
<p>Whichever row is furthest above its floor is where your engineering time should go. Practice on real data first by creating a free key on the signup page, then follow the Quickstart and test routes in the Sandbox.</p>
<ol>
<li>What to ask your data provider</li>
</ol>
<p>Armed with this model, you can ask sharper questions than "how fast are you?":</p>
<p>Where does your latency measurement start and end? (At event capture, at ingestion, or at emission?)
Is the figure a median or a tail? Ask for p95 and p99, not an average.
Which regions do you serve from, and which is closest to my users?
How are corrections represented? Are updates versioned, and is there a provisional/confirmed distinction?
What ordering guarantees exist per match?
What happens on reconnect? Is there replay, or must I fetch a snapshot?
What are the rate limits per plan, and what does the API return when I hit them? (See the pricing page.)
How do I verify webhook authenticity and handle retries? (See the Webhooks docs.)
13. Reference architecture and next steps
text
 Provider (REST + WebSocket + Webhooks)
            │
   ┌────────▼────────┐
   │  Ingest service │  parse → validate → dedupe → order per match
   └────────┬────────┘
            │ stamps trace context
   ┌────────▼────────┐
   │  State store    │  events keyed by id, derived score
   └────────┬────────┘
            │ pub/sub
   ┌────────▼────────┐
   │  Edge gateways  │  per region; conflate for slow clients
   └────────┬────────┘
            │ WebSocket / SSE
   ┌────────▼────────┐
   │     Clients     │  snapshot + deltas, rAF batching, resync on focus
   └─────────────────┘</p>
<p>Key takeaways:</p>
<p>Latency is a stack, not a number. Measure per layer.
The speed of light sets a floor. Put servers near users.
Fast and correct pull against each other. Model events and handle corrections.
Utilization near saturation makes latency explode. Keep headroom and conflate for slow clients.
Streaming beats polling mainly by removing repeated handshakes and interval staleness.
Track p95/p99, and remember that tails compound.</p>
<p>Want to build with this?</p>
<p>Read the documentation and the API reference.
Explore the Live Scores API, WebSocket API and Odds API.
Start free at Orbistats.</p>
<p>What's the biggest gap you've found between a provider's claimed latency and what your users actually experience? Share your layer-by-layer numbers in the comments. I'd love to compare notes.</p>
]]></content:encoded></item><item><title><![CDATA[Inside Orbistats' Research Section: What "Sub-50ms, End to End" Actually Requires]]></title><description><![CDATA[Introduction: Why "sub-50ms" deserves a second look
Every real-time data vendor has a number on its homepage. For Orbistats, the number is sub-50ms, sitting right next to 13 sports, REST, WebSocket, w]]></description><link>https://bettechmagnetics.hashnode.dev/inside-orbistats-research-section-what-sub-50ms-end-to-end-actually-requires</link><guid isPermaLink="true">https://bettechmagnetics.hashnode.dev/inside-orbistats-research-section-what-sub-50ms-end-to-end-actually-requires</guid><category><![CDATA[api]]></category><category><![CDATA[websockets]]></category><category><![CDATA[performance]]></category><category><![CDATA[backend]]></category><category><![CDATA[sports]]></category><dc:creator><![CDATA[Vijay Choudhary]]></dc:creator><pubDate>Mon, 05 Oct 2026 13:37:34 GMT</pubDate><enclosure url="https://cdn.hashnode.com/uploads/covers/6a47ac8dec0b8a17b4dc915d/073aeedd-e735-442a-bc4a-0950af54018f.png" length="0" type="image/jpeg"/><content:encoded><![CDATA[<p>Introduction: Why "sub-50ms" deserves a second look</p>
<p>Every real-time data vendor has a number on its homepage. For Orbistats, the number is sub-50ms, sitting right next to 13 sports, REST, WebSocket, webhooks, historical data and odds.</p>
<p>It's a good headline. But as a developer, the headline is where my questions begin:</p>
<p>Sub-50ms from what to what?
Is that a median, a p95, or a best case?
Measured where: inside a datacenter, or on my user's phone in Vadodara?
Does it include the time the real-world event took to reach the vendor?</p>
<p>Orbistats' Research section includes a piece titled "What sub-50ms actually requires, end to end." This post is my attempt to go deeper on the same question from the developer side: what does "end to end" really have to cover, and how do you verify any live feed yourself?</p>
<p>One honest note up front: the sub-50ms figure is the vendor's own marketing claim, not an independent benchmark. That isn't a criticism. It's the right starting point for any latency claim, and it's why the second half of this article is code you can run.</p>
<ol>
<li>"End to end" is a chain, not a number</li>
</ol>
<p>A goal is scored. Eventually a pixel changes on your user's screen. Between those two moments there are at least five hops:</p>
<p>text
Real-world event
      │
      ▼
[1] Source capture      (scout / official feed / data partner)
      │
      ▼
[2] Vendor ingest       (parse, validate, dedupe)
      │
      ▼
[3] Normalization       (one schema across sports and bookmakers)
      │
      ▼
[4] Fan-out / delivery  (WebSocket push, webhook POST, REST cache)
      │
      ▼
[5] Network + client    (TLS, routing, your server, your UI)
      │
      ▼
Pixel on screen</p>
<p>When a vendor says "sub-50ms", the honest follow-up is: which of these segments does it cover? The most common (and perfectly legitimate) interpretation is segments 3 → 4, from the moment the vendor has the event to the moment the packet leaves their edge. That is a real, valuable number. But it is not the same as "event to pixel."</p>
<p>An illustrative latency budget</p>
<p>This is not Orbistats' internal data. It's a generic model to show why the segments matter:</p>
<p>Segment	What happens	Typical order of magnitude
Source capture	Human scout or official data feed produces the event	Seconds (often dominates everything)
Vendor ingest	Parse, validate, dedupe	Single-digit ms
Normalization	Map to a unified schema	Single-digit ms
Fan-out	Push to connected subscribers	Single-digit to tens of ms
Network (vendor → you)	Distance, routing, TLS	~10–200+ ms depending on geography
Your app	Parse, render, repaint	Variable</p>
<p>Two takeaways:</p>
<p>The part a vendor controls is usually the smallest part. Source capture and your own network distance can dwarf it.
"Sub-50ms" and "the user saw it 3 seconds after the goal" can both be true. They're measuring different things.
2. What a vendor must actually get right to make a low-latency claim credible</p>
<p>Let me take the engineering view of each segment: what it requires, not what a brochure says.</p>
<p>2.1 Push, not poll</p>
<p>If you build live features on REST polling, your worst-case staleness is your polling interval, no matter how fast the server is. A 10-second poll means up to 10 seconds of lag, even from a 5ms API.</p>
<p>That's why WebSocket delivery and webhooks matter more for "live" than raw server speed:</p>
<p>text
REST polling:   client ──ask──▶ server ──answer──▶ client   (repeat every N seconds)
WebSocket:      client ◀═══ persistent connection ═══▶ server (server pushes on change)
Webhook:        server ──POST on event──▶ your endpoint</p>
<p>The arithmetic is also unforgiving. Polling all 13 sports once a minute is 13 × 1,440 = 18,720 requests/day. That exceeds typical free-tier limits many times over, and it still isn't real-time.</p>
<p>2.2 Persistent connections and efficient fan-out</p>
<p>Low latency at scale requires the vendor to avoid per-message connection setup, keep serialization cheap, and fan out one upstream event to many subscribers without re-processing it per client. In practice that means in-memory pub/sub, topic-based subscriptions (e.g. per sport or per match), and backpressure handling so one slow client can't stall the rest.</p>
<p>2.3 Normalization that doesn't sit on the hot path</p>
<p>Normalizing thirteen different sports, and many bookmakers' odds, into one schema is the product's value. Orbistats' Odds API describes normalized odds across bookmakers so you don't maintain a parser per bookmaker. For latency, the engineering requirement is that normalization is cheap and deterministic (table-driven mapping, no blocking I/O on the hot path).</p>
<p>2.4 Edge proximity</p>
<p>Speed-of-light is a hard floor. Cross-continent round trips alone can consume the entire 50ms budget, which is exactly why "measure from your own region" is the single most important piece of advice in this post.</p>
<p>2.5 Honest percentiles</p>
<p>A credible latency claim comes with p50 / p95 / p99, a measurement location, and a time window. Averages hide the spikes that users actually notice.</p>
<ol>
<li>Don't trust. Measure.</li>
</ol>
<p>Here's the practical part. These snippets let you evaluate any live feed, including Orbistats, from your own infrastructure.</p>
<p>Setup: Create a free key at orbistats.com/signup, then follow the Quickstart. The REST base URL is <a href="https://api.orbistats.com/v1/">https://api.orbistats.com/v1/</a> with Authorization: Bearer YOUR_API_KEY. Confirm the exact WebSocket URL and message shape in the documentation before running the streaming examples.</p>
<p>3.1 Baseline: REST round-trip time</p>
<p>First, measure plain HTTP round-trip from your region. This isolates network + server response time, not event-to-delivery lag.</p>
<p>js
// rtt.mjs  (Node 18+)
const BASE = "<a href="https://api.orbistats.com/v1">https://api.orbistats.com/v1</a>";
const KEY = process.env.ORBISTATS_KEY;</p>
<p>function percentile(sorted, p) {
  const idx = Math.min(sorted.length - 1, Math.ceil((p / 100) * sorted.length) - 1);
  return sorted[idx];
}</p>
<p>async function once(path) {
  const t0 = performance.now();
  const res = await fetch(<code>${BASE}${path}</code>, {
    headers: { Authorization: <code>Bearer ${KEY}</code> },
  });
  await res.arrayBuffer();               // include body download
  return performance.now() - t0;
}</p>
<p>const samples = [];
for (let i = 0; i &lt; 40; i++) {
  samples.push(await once("/football/matches/live"));
  await new Promise(r =&gt; setTimeout(r, 1500)); // be kind to rate limits
}</p>
<p>samples.sort((a, b) =&gt; a - b);
console.table({
  n: samples.length,
  p50: percentile(samples, 50).toFixed(1) + " ms",
  p95: percentile(samples, 95).toFixed(1) + " ms",
  p99: percentile(samples, 99).toFixed(1) + " ms",
  max: samples.at(-1).toFixed(1) + " ms",
});</p>
<p>Run it from your laptop, then from a cloud VM near your users, and compare. The difference is your geography tax.</p>
<p>⚠️ Mind your plan's request limits. Check the pricing page for current numbers, and keep sample counts small on the free tier.</p>
<p>3.2 Streaming: measure inter-message jitter and staleness</p>
<p>For a push feed, the most useful client-side metrics are:</p>
<p>Inter-arrival gaps (is the stream steady or bursty?)
Reconnect frequency (how often does it drop?)
Server-timestamp skew (if the payload carries an event timestamp)
js
// stream-probe.mjs
import WebSocket from "ws";</p>
<p>const WS_URL = process.env.ORBISTATS_WS_URL; // ← confirm in docs
const KEY = process.env.ORBISTATS_KEY;</p>
<p>let last = null;
const gaps = [];
const skews = [];</p>
<p>function connect(attempt = 0) {
  const ws = new WebSocket(WS_URL, {
    headers: { Authorization: <code>Bearer ${KEY}</code> },
  });</p>
<p>  ws.on("open", () =&gt; {
    console.log("connected");
    attempt = 0;
  });</p>
<p>  ws.on("message", (raw) =&gt; {
    const now = Date.now();
    const msg = JSON.parse(raw.toString());</p>
<pre><code>if (last) gaps.push(now - last);
last = now;

// If the payload carries an event/server timestamp (ms epoch), measure skew.
// Field name is an assumption: adapt to the real schema.
if (msg.ts) skews.push(now - msg.ts);
</code></pre>
<p>  });</p>
<p>  ws.on("close", () =&gt; {
    const delay = Math.min(30_000, 500 * 2 ** attempt) + Math.random() * 250;
    console.log(<code>closed; reconnecting in ${Math.round(delay)}ms</code>);
    setTimeout(() =&gt; connect(attempt + 1), delay);
  });</p>
<p>  ws.on("error", (e) =&gt; console.error("ws error:", e.message));
}</p>
<p>connect();</p>
<p>setInterval(() =&gt; {
  if (!gaps.length) return;
  const g = [...gaps].sort((a, b) =&gt; a - b);
  console.log("gap p50/p95:", g[Math.floor(g.length * 0.5)], g[Math.floor(g.length * 0.95)], "ms");
  if (skews.length) {
    const s = [...skews].sort((a, b) =&gt; a - b);
    console.log("skew p50/p95:", s[Math.floor(s.length * 0.5)], s[Math.floor(s.length * 0.95)], "ms");
  }
}, 30_000);</p>
<p>A caveat that matters: client-side skew only means something if your clock is synced (use NTP/chrony) and the payload exposes a trustworthy timestamp. If it doesn't, rely on gap and reconnect statistics, and don't invent a latency number you can't justify.</p>
<p>3.3 The architecture that keeps your key (and your latency) safe</p>
<p>Never call the API from the browser with your key. Put a thin relay in the middle:</p>
<p>text
Orbistats ──WS──▶ Your server ──WS/SSE──▶ Your users' browsers
                     │
                     └── cache + fan-out + key stays server-side</p>
<p>One upstream connection, many downstream clients. That protects your key and your rate limit, and it often improves perceived latency because your users connect to a server you can place near them.</p>
<ol>
<li>Choosing the right delivery method for your latency needs
Need	Best fit	Why
Live scoreboard on a site	WebSocket API	Push, low overhead, one connection
Trigger actions (notifications, settlement)	Webhooks	Event-driven; no polling loop
Match page load / initial state	REST via Live Scores API	Simple snapshot, cacheable
Odds boards and line movement	Odds API + WebSocket	Prices change fast; push beats poll
Fast UI with minimal code	Widgets	Drop-in components
Model training and backtests	Historical Sports Data API	Latency irrelevant; coverage matters</li>
</ol>
<p>Notice what's not in this table: a recommendation to chase the lowest millisecond. For most products, correctness, coverage and reliability beat shaving 20ms.</p>
<ol>
<li>When latency really matters (and when it doesn't)</li>
</ol>
<p>It matters a lot when you're building trading tools, in-play pricing displays, or anything where seconds of staleness changes a decision. Even then, the dominant lag is usually upstream of the vendor (how the event was captured), so ask vendors where their timestamps come from.</p>
<p>It matters much less for fixtures and standings, season-long analytics, content sites, fantasy products (where scores update on a minute scale), and historical research.</p>
<p>If you're not sure which camp you're in, you probably don't need the lowest-latency tier. Start free, measure, and upgrade when data proves the need. Free-tier live data may be delayed relative to paid plans, so check the current plan details on the pricing page.</p>
<ol>
<li>A checklist to evaluate any "sub-X ms" claim</li>
</ol>
<p>Copy this into your next vendor evaluation:</p>
<p>✅ What's the start and end point? Event capture → pixel, or vendor ingest → vendor edge?
✅ Which percentile? p50 alone is marketing; ask for p95/p99.
✅ Measured where? Same datacenter, or a user region like yours?
✅ Push or poll? If poll, your interval is your floor.
✅ Is there a status page? Look for something like Orbistats' status page and a public changelog.
✅ Is there a sandbox? Test before you pay: Sandbox.
✅ Is the schema stable? Versioning (/v1/) and a documented change policy protect your integration.
✅ Can you reproduce the number yourself? If not, treat it as a hypothesis.
7. One API, 13 sports: the multiplier effect</p>
<p>Latency is one axis. Breadth is the other. A single integration pattern across thirteen sports (Football, Basketball, American Football, Cricket, Tennis, Baseball, Esports, Combat Sports, Volleyball, Handball, Ice Hockey, Golf and Horse Racing) means you swap a path segment instead of writing a new integration:</p>
<p>js
const SPORTS = [
  "football", "basketball", "american-football", "cricket", "tennis",
  "baseball", "esports", "combat-sports", "volleyball", "handball",
  "ice-hockey", "golf", "horse-racing",
];</p>
<p>async function liveFor(sport) {
  const res = await fetch(<code>https://api.orbistats.com/v1/${sport}/matches/live</code>, {
    headers: { Authorization: <code>Bearer ${process.env.ORBISTATS_KEY}</code> },
  });
  if (!res.ok) throw new Error(<code>${sport}: ${res.status}</code>);
  return res.json();
}</p>
<p>Sports pages like Football document what's covered per sport. One caution: a shared schema doesn't make the sports identical. Odds semantics, event types and "what counts as live" differ between, say, tennis and horse racing, so test each sport you ship.</p>
<ol>
<li>Production habits that protect your latency</li>
</ol>
<p>A few things that quietly make "real-time" apps feel slow:</p>
<p>Cache aggressively for slow-changing resources (fixtures, standings, teams). Spend your request budget on live data.
Reconnect with exponential backoff and jitter, as in the probe above, so a vendor blip doesn't become a thundering herd.
Heartbeat and stale-detection: if you haven't received a message in N seconds during a live match, show a "reconnecting" state instead of silently displaying old data.
Deduplicate on match_id + event sequence so redelivery after reconnect doesn't double-count a goal.
Keep keys server-side, always.</p>
<p>For the full reference of resources and parameters, use the API Reference and the Developer Hub.</p>
<p>Conclusion</p>
<p>"Sub-50ms, end to end" is only as meaningful as its definition. The honest version of the claim names the segments it covers, the percentile it reports, and the location it's measured from, and then invites you to verify it.</p>
<p>That's the real lesson from reading a vendor's latency research with an engineer's eyes: the number is where the conversation starts, not where it ends. Push delivery beats polling. Geography beats micro-optimization. Percentiles beat averages. And a ten-minute probe from your own region beats any homepage.</p>
<p>If you want to run the experiments above, you can create a free API key, explore the Orbistats homepage, or read the original Research section.</p>
<p>Your turn: what latency numbers do you see from your region? Drop your p50/p95 in the comments. I'd love to compare notes across geographies.</p>
<p>Disclaimer: Orbistats' latency figures are the vendor's own claims. Odds data is informational only, not betting advice.</p>
]]></content:encoded></item><item><title><![CDATA[Sportsbooks vs Trading Desks vs Media: How Orbistats' Solution Pages Map to Real Architectures]]></title><description><![CDATA[Browse the Solutions section of Orbistats and you will notice something that looks like marketing but is actually engineering. Sportsbooks, trading desks and media publishers each get their own page, ]]></description><link>https://bettechmagnetics.hashnode.dev/sportsbooks-vs-trading-desks-vs-media-how-orbistats-solution-pages-map-to-real-architectures</link><guid isPermaLink="true">https://bettechmagnetics.hashnode.dev/sportsbooks-vs-trading-desks-vs-media-how-orbistats-solution-pages-map-to-real-architectures</guid><category><![CDATA[System Design]]></category><category><![CDATA[sports data]]></category><category><![CDATA[API Design]]></category><category><![CDATA[websocket]]></category><category><![CDATA[backend]]></category><dc:creator><![CDATA[Vijay Choudhary]]></dc:creator><pubDate>Mon, 05 Oct 2026 13:34:26 GMT</pubDate><enclosure url="https://cdn.hashnode.com/uploads/covers/6a47ac8dec0b8a17b4dc915d/5596fa08-b016-4c54-8b42-2e7e248472ef.png" length="0" type="image/jpeg"/><content:encoded><![CDATA[<p>Browse the Solutions section of Orbistats and you will notice something that looks like marketing but is actually engineering. Sportsbooks, trading desks and media publishers each get their own page, with different headlines and different feature emphasis.</p>
<p>The underlying API is the same. The systems built on top of it are completely different.</p>
<p>A sportsbook cares about market coverage and correct settlement. A trading desk cares about milliseconds and line movement. A media site cares about readable live pages that survive a traffic spike. The same data, three architectures, three sets of trade-offs.</p>
<p>This post maps each solution page to the architecture it implies, with diagrams and code you can adapt. By the end you should know which Orbistats products to reach for, how to wire them together, and which mistakes to avoid.</p>
<p>Note: The code here is illustrative. For exact endpoints, payloads and limits, use the official documentation as the source of truth.</p>
<p>TL;DR
Sportsbooks need breadth: many markets, correct results, and reliable settlement. Think pull-heavy, correctness-first. See Sportsbooks &amp; Trading.
Trading desks need speed and change detection: persistent connections, in-memory state, line movement signals. Think push-based, latency-first. See Trading Desks.
Media and publishers need presentation and resilience: live scores, match centers, widgets, caching. Think read-heavy, availability-first. See Media &amp; Publishers.
All three share the same foundation: one normalized schema, a REST API, WebSocket streaming and webhooks.
Pick your delivery method by workload, not habit: REST for snapshots, WebSocket for continuous streams, webhooks for event-driven actions.
Table of Contents
One API, many architectures
The shared foundation
Choosing a delivery method (REST vs WebSocket vs Webhooks)
Architecture 1: The sportsbook
Architecture 2: The trading desk
Architecture 3: The media and publisher site
Side-by-side comparison
The other four solution pages (Fantasy, Analytics, Teams, Leagues)
13 sports: what changes per architecture
Cross-cutting concerns: auth, rate limits, retries, versioning
Testing your architecture
Common mistakes
Which architecture are you building? A decision guide
FAQ</p>
<ol>
<li>One API, many architectures</li>
</ol>
<p>Orbistats is positioned as a sports data and odds API covering 13 sports. The product surface includes:</p>
<p>Sports Data API: fixtures, results, standings, teams, competitions
Live Scores API: real-time match events
Odds API: normalized odds, markets and line movement
Historical Sports Data API: multi-season archive
WebSocket API and Webhooks: push delivery
Widgets: drop-in UI</p>
<p>Different customers combine these pieces in different ways. The Solutions pages exist because "what should I use?" depends entirely on what you are building. Reading them as architecture guides, rather than sales pages, is a good way to save weeks of design time.</p>
<ol>
<li>The shared foundation</li>
</ol>
<p>Before we split into three architectures, here is what they all have in common.</p>
<p>text
                    ORBISTATS
                        │
        ┌───────────────┼───────────────┐
        ▼               ▼               ▼
   Sports data      Odds API       Historical
   (fixtures,       (markets,      (archive,
    results,         lines,         closing odds)
    stats)           movement)
        │               │               │
        └───────────────┼───────────────┘
                        ▼
              One normalized schema
                        │
        ┌───────────────┼───────────────┐
        ▼               ▼               ▼
       REST         WebSocket        Webhooks</p>
<p>Every integration starts the same way:</p>
<p>http
GET <a href="https://api.orbistats.com/v1/fixtures">https://api.orbistats.com/v1/fixtures</a>
Authorization: Bearer YOUR_API_KEY</p>
<p>The base URL is <a href="https://api.orbistats.com/v1/">https://api.orbistats.com/v1/</a> and authentication uses a bearer token. Because the schema is normalized, the structural code you write for one sport or one data type transfers to the others. If you want the details of how that normalization works, we covered it in our post on odds normalization.</p>
<p>Here is a small shared client you can reuse in all three architectures:</p>
<p>python
import os
import time
import requests
from requests.adapters import HTTPAdapter
from urllib3.util.retry import Retry</p>
<p>BASE_URL = "<a href="https://api.orbistats.com/v1">https://api.orbistats.com/v1</a>"</p>
<p>def make_session(api_key: str) -&gt; requests.Session:
    s = requests.Session()
    s.headers.update({"Authorization": f"Bearer {api_key}"})
    retry = Retry(
        total=3,
        backoff_factor=0.5,
        status_forcelist=(429, 500, 502, 503, 504),
        allowed_methods=("GET",),
    )
    s.mount("https://", HTTPAdapter(max_retries=retry, pool_maxsize=20))
    return s</p>
<p>class Orbistats:
    def <strong>init</strong>(self, api_key: str | None = None):
        self.session = make_session(api_key or os.environ["ORBISTATS_API_KEY"])</p>
<pre><code>def get(self, path: str, **params):
    r = self.session.get(f"{BASE_URL}{path}", params=params, timeout=10)
    r.raise_for_status()
    return r.json()
</code></pre>
<p>Connection pooling, bounded retries with backoff and an explicit timeout. These are table stakes in all three architectures.</p>
<ol>
<li>Choosing a delivery method (REST vs WebSocket vs Webhooks)</li>
</ol>
<p>This choice shapes everything else, so get it right first.</p>
<pre><code>REST	WebSocket	Webhooks
</code></pre>
<p>Direction	You ask, server answers	Persistent, server pushes	Server calls your URL
Best for	Snapshots, backfills, page loads	Continuous live streams	Event-driven reactions
Latency to change	Depends on poll interval	Lowest	Low, per event
Cost at scale	Grows with polling frequency	Constant per connection	Grows with event volume
Your infra	Stateless	Stateful (connection, buffers)	Public endpoint, queue
Failure mode	Rate limits	Disconnects, backpressure	Lost or duplicate deliveries</p>
<p>A rule of thumb: if you are polling the same endpoint more than once every few seconds, you probably want WebSocket or webhooks instead. Polling live data is both slower and more expensive.</p>
<ol>
<li>Architecture 1: The sportsbook
What the Solutions page emphasizes</li>
</ol>
<p>Sportsbooks &amp; Trading leads with odds, markets and results. That is a clue to the real needs: wide market coverage, trustworthy settlement data and a reference price to benchmark against.</p>
<p>The core job</p>
<p>A sportsbook is not a data consumer in the same way a media site is. It makes its own prices, then needs an external reference to sanity-check them, plus authoritative results to settle bets. Orbistats typically sits in three places in that picture:</p>
<p>Reference odds to compare against your own lines
Fixtures and schedule to create the events you offer
Results to settle markets
Architecture
text
┌──────────────┐     ┌────────────────────────┐
│ Orbistats    │────▶│ Ingestion workers      │
│ REST + WS    │     │ (fixtures, odds, results)
└──────────────┘     └───────────┬────────────┘
                                 ▼
                     ┌────────────────────────┐
                     │ Internal market store  │  (DB, source of truth)
                     └───────────┬────────────┘
        ┌────────────────────────┼─────────────────────┐
        ▼                        ▼                     ▼
┌───────────────┐       ┌────────────────┐    ┌─────────────────┐
│ Pricing /     │       │ Bet acceptance │    │ Settlement      │
│ risk engine   │       │ service        │    │ service         │
│ (your margin) │       │                │    │ (results-driven)│
└───────┬───────┘       └────────┬───────┘    └────────┬────────┘
        └────────────────────────┼─────────────────────┘
                                 ▼
                       Customer-facing product
Pattern A: benchmark your prices against the market</p>
<p>The simplest high-value use of external odds is a price deviation monitor. If your line is far from the market, you either have an opportunity or a mistake.</p>
<p>python
def devig_proportional(decimals: list[float]) -&gt; list[float]:
    probs = [1 / d for d in decimals]
    total = sum(probs)
    return [p / total for p in probs]</p>
<p>def deviation_report(our_odds: dict, market_odds: dict, threshold: float = 0.03):
    """
    our_odds / market_odds: {"home": 1.95, "draw": 3.4, "away": 4.1}
    Compares fair probabilities and flags outcomes that drift too far.
    """
    keys = list(our_odds.keys())
    ours = devig_proportional([our_odds[k] for k in keys])
    mkt = devig_proportional([market_odds[k] for k in keys])</p>
<pre><code>flags = []
for k, p_ours, p_mkt in zip(keys, ours, mkt):
    diff = p_ours - p_mkt
    if abs(diff) &gt;= threshold:
        flags.append({"outcome": k, "our_prob": round(p_ours, 4),
                      "market_prob": round(p_mkt, 4), "diff": round(diff, 4)})
return flags
</code></pre>
<p>Comparing fair probabilities (margin removed) instead of raw prices is important, because your margin and the market's margin will differ by design.</p>
<p>Pattern B: settlement driven by results</p>
<p>Settlement is where correctness matters most. A late or wrong result is a financial problem. Design settlement as an idempotent, auditable job:</p>
<p>python
from enum import Enum</p>
<p>class Outcome(str, Enum):
    WIN = "win"
    LOSE = "lose"
    VOID = "void"
    PUSH = "push"</p>
<p>def settle_total(line: float, home: int, away: int, side: str) -&gt; Outcome:
    goals = home + away
    if goals == line:
        return Outcome.PUSH
    if side == "over":
        return Outcome.WIN if goals &gt; line else Outcome.LOSE
    return Outcome.WIN if goals &lt; line else Outcome.LOSE</p>
<p>def settle_event(db, client, event_id: str):
    # 1. Idempotency: never settle twice
    if db.is_settled(event_id):
        return</p>
<pre><code>result = client.get(f"/football/results/{event_id}")
if result.get("status") != "finished":
    return                                # only settle finished events

home, away = result["home"]["score"], result["away"]["score"]

with db.transaction():
    for bet in db.open_bets(event_id):
        outcome = settle_total(bet.line, home, away, bet.side)
        db.record_settlement(bet.id, outcome, source_snapshot=result)
    db.mark_settled(event_id)
</code></pre>
<p>Three practices to copy:</p>
<p>Store the source snapshot with each settlement so you can audit it later.
Only settle on a final status, never on a live score.
Handle void and push explicitly. Postponements, abandoned matches and retirements need rules, and those rules vary by sport. See the per-sport table later in this post.
Pattern C: schedule sync</p>
<p>Events must exist before you can offer markets. A periodic sync of fixtures keeps your catalog fresh:</p>
<p>python
def sync_fixtures(db, client, sport: str, competition: str):
    data = client.get("/fixtures", sport=sport, competition=competition)
    for fx in data.get("data", []):
        db.upsert_event(
            external_id=fx["match_id"],
            start_time=fx["start_time"],
            home=fx["home"]["name"],
            away=fx["away"]["name"],
            status=fx["status"],
        )</p>
<p>Run it on a schedule, store the external ID from Orbistats as a mapping key, and treat reschedules (a changed start_time) as first-class events that trigger market suspension and customer notification.</p>
<p>Why this is "pull-heavy"</p>
<p>Most sportsbook ingestion is scheduled and batch-like. WebSocket becomes relevant mainly for in-play products, where you need fast score and odds changes. Pre-match can run happily on REST.</p>
<ol>
<li>Architecture 2: The trading desk
What the Solutions page emphasizes</li>
</ol>
<p>Trading Desks leads with low-latency odds, WebSocket and line movement. That is a completely different design brief.</p>
<p>The core job</p>
<p>A trading desk reacts to change. The question is not "what is the price?" but "what just moved, by how much, and across how many bookmakers?" The architecture is built around a persistent stream, an in-memory state and fast decision logic.</p>
<p>The homepage describes a sub-50ms live feed and Orbistats publishes research on what sub-50ms actually requires end to end. Whatever your own latency budget is, the principle holds: the feed is only one part of the latency chain. Your network, parsing, queuing and decision code are the rest.</p>
<p>Architecture
text
┌──────────────┐  WebSocket   ┌──────────────────────┐
│ Orbistats    │─────────────▶│ Stream consumer      │
│ odds stream  │              │ (single reader/conn) │
└──────────────┘              └──────────┬───────────┘
                                         ▼
                              ┌──────────────────────┐
                              │ In-memory order book │  latest price per
                              │ (event/market/book)  │  (event, market, bookmaker)
                              └──────────┬───────────┘
                    ┌────────────────────┼───────────────────┐
                    ▼                    ▼                   ▼
          ┌────────────────┐  ┌──────────────────┐  ┌────────────────┐
          │ Movement       │  │ Best-price /     │  │ Persistence    │
          │ detector       │  │ consensus calc   │  │ (async, batched)│
          └───────┬────────┘  └────────┬─────────┘  └────────────────┘
                  └────────────────────┼──────────────────┘
                                       ▼
                              Signals / alerts / orders</p>
<p>Notice what is not on the hot path: the database. Writes happen asynchronously and in batches, so storage latency never blocks the decision loop.</p>
<p>Pattern A: a resilient WebSocket consumer</p>
<p>A trading consumer must survive disconnects without losing state or hammering the server.</p>
<p>python
import asyncio
import json
import random
import websockets</p>
<p>WS_URL = "wss://api.orbistats.com/v1/stream"   # illustrative: check docs for the exact URL</p>
<p>async def consume(api_key: str, handle_update):
    attempt = 0
    while True:
        try:
            async with websockets.connect(
                f"{WS_URL}?token={api_key}",
                ping_interval=20,
                ping_timeout=20,
                max_queue=1024,
            ) as ws:
                attempt = 0                      # reset after a good connect
                await ws.send(json.dumps({
                    "action": "subscribe",
                    "channel": "odds",
                    "sport": "football",
                }))
                async for message in ws:
                    await handle_update(json.loads(message))
        except (websockets.ConnectionClosed, OSError):
            attempt += 1
            # exponential backoff with jitter, capped
            delay = min(30, 2 ** attempt) * (0.5 + random.random() / 2)
            await asyncio.sleep(delay)
            # NOTE: after reconnect, re-sync a snapshot via REST so you
            # don't trade on state you missed while disconnected.</p>
<p>The most important comment is the last one. After any reconnect, rebuild your state from a REST snapshot, because you missed updates while you were away. A stream alone is not enough. The robust pattern is snapshot plus deltas.</p>
<p>Pattern B: in-memory book and movement detection
python
import time
from collections import defaultdict, deque</p>
<p>class OddsBook:
    def <strong>init</strong>(self, history: int = 50):
        # key: (event_id, market, line, outcome, bookmaker) -&gt; deque[(ts, decimal)]
        self.prices = defaultdict(lambda: deque(maxlen=history))</p>
<pre><code>def update(self, key, decimal: float, ts: float | None = None):
    ts = ts or time.time()
    dq = self.prices[key]
    prev = dq[-1][1] if dq else None
    dq.append((ts, decimal))
    return prev

def latest(self, key):
    dq = self.prices.get(key)
    return dq[-1][1] if dq else None
</code></pre>
<p>def movement(prev: float | None, new: float) -&gt; dict | None:
    if prev is None or prev == new:
        return None
    # Work in implied probability so moves are comparable across price levels
    p_prev, p_new = 1 / prev, 1 / new
    return {
        "from": prev, "to": new,
        "prob_delta": round(p_new - p_prev, 5),
        "direction": "shortened" if new &lt; prev else "drifted",
    }</p>
<p>book = OddsBook()</p>
<p>async def handle_update(msg):
    key = (msg["event_id"], msg["market"], msg.get("line"),
           msg["outcome"], msg["bookmaker"])
    prev = book.update(key, msg["decimal"])
    mv = movement(prev, msg["decimal"])
    if mv and abs(mv["prob_delta"]) &gt;= 0.02:
        await emit_signal(msg["event_id"], msg["market"], key, mv)</p>
<p>Measuring movement in implied probability instead of raw price differences matters. A move from 1.50 to 1.45 and a move from 10.0 to 9.5 are very different in probability terms. Raw price deltas hide that.</p>
<p>Pattern C: cross-bookmaker consensus</p>
<p>A single bookmaker moving alone is often noise. Many moving together is information. A simple consensus check:</p>
<p>python
from statistics import median</p>
<p>def consensus_move(book: OddsBook, event_id, market, line, outcome,
                   bookmakers, window_s: float = 30.0, min_books: int = 3):
    now = time.time()
    shortening = 0
    for b in bookmakers:
        dq = book.prices.get((event_id, market, line, outcome, b))
        if not dq or len(dq) &lt; 2:
            continue
        recent = [p for (t, p) in dq if now - t &lt;= window_s]
        if len(recent) &gt;= 2 and recent[-1] &lt; recent[0]:
            shortening += 1
    return shortening &gt;= min_books</p>
<p>This is where normalization pays off. The loop treats every bookmaker identically because the data arrives in one schema.</p>
<p>Backpressure: the unglamorous killer</p>
<p>If updates arrive faster than you process them, memory grows and latency climbs silently. Use a bounded queue and decide your policy explicitly:</p>
<p>python
import asyncio</p>
<p>queue: asyncio.Queue = asyncio.Queue(maxsize=10_000)</p>
<p>async def producer(msg):
    try:
        queue.put_nowait(msg)
    except asyncio.QueueFull:
        # Policy choice: drop the oldest (we only care about latest price)
        _ = queue.get_nowait()
        queue.put_nowait(msg)
        metrics.incr("odds.dropped_stale")</p>
<p>For latest-price state, dropping older updates under load is usually correct, because the newest price supersedes the old one. Whatever policy you choose, make it a decision and measure it.</p>
<p>Webhooks for the slow lane</p>
<p>Not every trading action needs a stream. Alerts, notifications and workflow triggers can use webhooks, which keeps long-lived connections for the hot path only.</p>
<ol>
<li>Architecture 3: The media and publisher site
What the Solutions page emphasizes</li>
</ol>
<p>Media &amp; Publishers leads with live scores, match centers and widgets. The brief is totally different again: look good, load fast, and survive traffic spikes.</p>
<p>The core job</p>
<p>A publisher wants a live match page that hundreds of thousands of readers can open during a big game without knocking over the site. The design question is how to serve the same data to many readers cheaply.</p>
<p>The answer is a classic: fetch once, cache, fan out. You never let reader traffic reach the API directly.</p>
<p>Architecture
text
┌──────────────┐        ┌───────────────────────┐
│ Orbistats    │───────▶│ Poller / WS consumer  │  1 consumer, not N readers
│ Live Scores  │        └───────────┬───────────┘
└──────────────┘                    ▼
                          ┌───────────────────┐
         Webhook ────────▶│ Cache (Redis/KV)  │
      (event invalidates) └─────────┬─────────┘
                                    ▼
                          ┌───────────────────┐
                          │ Your web tier /   │
                          │ CDN / edge        │
                          └─────────┬─────────┘
                                    ▼
                         Thousands of readers
Option 1: the fast path, drop-in widgets</p>
<p>If you do not want to build a UI, Widgets cover common publisher needs such as live scores, match centers and odds boards, embedded into your pages. This is the lowest-effort route and a good fit for editorial teams without much front-end capacity. Check the widgets page for the current embed method and customization options.</p>
<p>Option 2: the build-your-own path</p>
<p>When you want full control over design, build a thin backend that owns the API relationship and exposes a cheap read endpoint to your front end.</p>
<p>python
import json
import time
import redis
from fastapi import FastAPI, Response</p>
<p>r = redis.Redis(host="localhost", port=6379, decode_responses=True)
app = FastAPI()</p>
<p>LIVE_TTL = 5          # seconds: short for live matches
FIXTURE_TTL = 300     # seconds: longer for schedule data</p>
<p>def cache_key(sport: str, match_id: str) -&gt; str:
    return f"match:{sport}:{match_id}"</p>
<h1>---- Reader-facing endpoint: NEVER calls Orbistats directly ----</h1>
<p>@app.get("/api/match/{sport}/{match_id}")
def read_match(sport: str, match_id: str, response: Response):
    cached = r.get(cache_key(sport, match_id))
    if cached is None:
        response.status_code = 404
        return {"error": "not_available"}
    response.headers["Cache-Control"] = "public, max-age=3, stale-while-revalidate=10"
    return json.loads(cached)</p>
<h1>---- Background refresher: the ONLY thing that calls Orbistats ----</h1>
<p>def refresh_live(client, sport: str):
    data = client.get(f"/{sport}/matches/live")
    pipe = r.pipeline()
    for m in data.get("data", []):
        pipe.setex(cache_key(sport, m["match_id"]), LIVE_TTL * 3, json.dumps(m))
    pipe.execute()</p>
<p>This design gives you a powerful property: your API usage is independent of your audience size. Whether 100 or 1,000,000 people are watching, the refresher makes the same number of calls. That is how you stay inside rate limits and keep costs predictable. See pricing for plan limits.</p>
<p>Webhooks to refresh the cache on events</p>
<p>Instead of polling on a timer, let events drive refreshes. A goal triggers a webhook, your handler refreshes that one match, and the cache stays fresh with fewer calls.</p>
<p>python
@app.post("/webhook/sports")
async def sports_hook(payload: dict):
    if payload.get("type") in {"goal", "score_update", "status_change"}:
        sport = payload["sport"]
        match_id = payload["match_id"]
        fresh = client.get(f"/{sport}/matches/{match_id}")
        r.setex(cache_key(sport, match_id), LIVE_TTL * 3, json.dumps(fresh))
    return {"ok": True}</p>
<p>Return 2xx fast, do work on a queue if it is heavy, and make the handler idempotent, because deliveries can repeat.</p>
<p>SEO and rendering considerations</p>
<p>Publishers care about search traffic, so a few notes:</p>
<p>Server-side render the initial match page so crawlers and slow devices get content immediately, then hydrate live updates on the client.
Use structured data (SportsEvent schema) on match pages.
For live updates in the browser, poll your own cached endpoint (cheap) or use server-sent events from your own backend. Do not put your API key in front-end code.
javascript
// Client-side: polls YOUR cached endpoint, never Orbistats directly
async function refreshScore(sport, matchId) {
  const res = await fetch(<code>/api/match/${sport}/${matchId}</code>);
  if (!res.ok) return;
  const m = await res.json();
  document.querySelector("#score").textContent =
    <code>${m.home.name} ${m.home.score} - ${m.away.score} ${m.away.name}</code>;
}
setInterval(() =&gt; refreshScore("football", "match_50231"), 5000);
Combining odds with editorial</p>
<p>Publishers often add an odds board or "odds movement" box to match previews. Pull it through the same cache pattern, and read the glossary so your editorial copy explains terms like implied probability correctly.</p>
<ol>
<li>Side-by-side comparison
Dimension	Sportsbook	Trading desk	Media / publisher
Primary goal	Correct markets and settlement	React to change fast	Serve many readers cheaply
Main data	Fixtures, results, reference odds	Live odds, line movement	Live scores, events, match data
Delivery method	REST (+ WS for in-play)	WebSocket-first	REST + cache, webhooks
Latency sensitivity	Medium	Very high	Low to medium
State	Database is source of truth	In-memory book	Cache (Redis/CDN)
Failure to avoid	Wrong settlement	Stale or missed updates	Origin overload
Scaling lever	Idempotent batch jobs	Backpressure and sharding	Caching and fan-out
Typical plan fit	Scales with market coverage	Needs real-time tier	Scales with traffic via cache
Solution page	Sportsbooks	Trading	Media</li>
</ol>
<p>The surprising takeaway: the sportsbook is the least latency-sensitive of the three. Its hard problem is correctness, not speed.</p>
<ol>
<li>The other four solution pages</li>
</ol>
<p>The same logic maps to the remaining Solutions pages.</p>
<p>Fantasy &amp; AI Data</p>
<p>Fantasy &amp; AI Data emphasizes player data and historical data. The architecture is a feature pipeline: pull player and match statistics, store them in a warehouse, compute features such as rolling form and minutes played, then score players for contests. The Sports Statistics API is the natural source. Latency matters mainly at lineup lock, so schedule heavy syncs ahead of it.</p>
<p>Analytics &amp; Data Science</p>
<p>Analytics &amp; Data Science emphasizes structured historical datasets and bulk exports. The architecture is a batch ingestion job into a data lake or warehouse, followed by notebooks and model training. Use the Historical Sports Data API for backtesting, and version your datasets so experiments are reproducible.</p>
<p>python
import pandas as pd</p>
<p>def load_season(client, sport: str, season: int) -&gt; pd.DataFrame:
    rows, page = [], 1
    while True:
        resp = client.get("/results", sport=sport, season=season, page=page)
        batch = resp.get("data", [])
        if not batch:
            break
        rows.extend(batch)
        page += 1
    df = pd.json_normalize(rows)
    df.to_parquet(f"data/{sport}_{season}.parquet")   # snapshot for reproducibility
    return df
Teams &amp; Clubs</p>
<p>Teams &amp; Clubs emphasizes performance and player/team history. The architecture is an internal performance dashboard: historical trends, opponent scouting and player workload, refreshed after each match.</p>
<p>Federations &amp; Leagues</p>
<p>Federations &amp; Leagues emphasizes fixtures, standings and custom feeds. The architecture is a publishing and distribution layer: official schedules and tables feeding the league's own website, apps and partner feeds.</p>
<ol>
<li>13 sports: what changes per architecture</li>
</ol>
<p>Orbistats covers 13 sports: Football, Basketball, American Football, Cricket, Tennis, Baseball, Esports, Combat Sports, Volleyball, Handball, Ice Hockey, Golf and Horse Racing. Each architecture feels sport differences in a different place.</p>
<p>Sport	Sportsbook impact	Trading impact	Media impact
Football	3-way markets, regulation vs ET	High in-play volume	Minute-by-minute match center
Basketball	Overtime in moneyline, many alt lines	Very frequent price changes	Quarter-based scoreboards
American Football	Key numbers in spreads	Sharp moves around news	Drive and down tracking
Cricket	Format-specific rules (T20, ODI, Test)	Over-by-over repricing	Overs and innings display
Tennis	Retirement and walkover rules	Point-by-point live trading	Nested set and game scores
Baseball	Pitcher-dependent markets	Lineup news moves lines	Inning-based boxscores
Esports	Series vs map markets	Fast, volatile markets	Series and map views
Combat Sports	Method and round markets	Sparse, event-driven moves	Fight card pages
Volleyball	Set and point handicaps	Set-level repricing	Set scoreboards
Handball	High-scoring totals	Fast total shifts	Half-based scores
Ice Hockey	Regulation vs OT moneyline	Goal-driven jumps	Period scoreboards
Golf	Outrights, each-way, dead heats	Slow-moving outright markets	Leaderboards
Horse Racing	Each-way terms, non-runners	Pre-race price drifts	Racecards and results</p>
<p>The pattern: sportsbooks feel sports differences in settlement rules, trading desks feel them in update frequency, and media sites feel them in how the scoreboard is shaped. Keep sport-specific logic in an adapter layer so your core architecture stays sport-agnostic.</p>
<ol>
<li>Cross-cutting concerns</li>
</ol>
<p>These apply to all three architectures.</p>
<p>Authentication. Use the bearer token from a backend only. Never ship your key to browsers or mobile apps. Proxy through your own service.</p>
<p>Rate limits and plans. Free is for testing, and paid tiers add real-time production use, historical access, widgets and webhooks. Design your polling and caching to fit your plan. See pricing.</p>
<p>Retries. Retry idempotent GETs with exponential backoff and jitter. Never retry in a tight loop.</p>
<p>Versioning. The API is versioned (/v1/ today). Put the version in one config constant and keep your parsing tolerant of new fields, so non-breaking additions never break you.</p>
<p>Observability. Track request latency, error rates, stream reconnects, queue depth, cache hit rate and "last update age" per event. The last one is the best single indicator of whether your data is alive.</p>
<p>python
def freshness_seconds(updated_at_iso: str) -&gt; float:
    from datetime import datetime, timezone
    ts = datetime.fromisoformat(updated_at_iso.replace("Z", "+00:00"))
    return (datetime.now(timezone.utc) - ts).total_seconds()</p>
<p>Status awareness. Link your alerts to the status page and watch the changelog.</p>
<p>Compliance. Betting-related products are regulated in many jurisdictions. Check licensing, age-gating, geo-restrictions and responsible-gambling requirements for your market, and review Orbistats' terms and data licensing before you ship.</p>
<ol>
<li>Testing your architecture</li>
</ol>
<p>Each architecture has a different test focus.</p>
<p>Sportsbook: unit-test settlement for every sport and edge case (push, void, extra time). Replay historical results through the settlement code and compare outcomes.
Trading: replay recorded streams through your handler at faster-than-real-time to find bottlenecks. Inject disconnects to verify snapshot-then-delta recovery.
Media: load-test your cached endpoint, not the upstream API. Verify that a cache miss never stampedes the origin.</p>
<p>A quick way to start without writing integration scaffolding is the Sandbox, plus the Quickstart and Examples.</p>
<ol>
<li>Common mistakes
Using one architecture for everything. A media-style cache is wrong for a trading desk, and a trading-style stream is overkill for a results page.
Polling live data. Use WebSocket or webhooks for anything that changes every few seconds.
Putting the database on the trading hot path. Persist asynchronously.
No snapshot after reconnect. A stream with gaps is worse than no stream, because you will not know you are wrong.
Letting readers hit the API. For media, readers should hit your cache, always.
Settling on live or non-final data. Settle only on final statuses, and make it idempotent.
Ignoring sport-specific rules. Retirement, overtime and dead-heat rules are where money is lost.
Exposing API keys client-side. Proxy through your backend.
Comparing raw prices across bookmakers. Compare fair probabilities after removing margin.
No freshness metric. If you cannot see data age, you will serve stale data without knowing it.</li>
<li>Which architecture are you building? A decision guide</li>
</ol>
<p>Answer these in order:</p>
<p>Do you set prices and take bets? Start from the sportsbook pattern: reference odds, fixtures and results.
Do you act on price changes within seconds? Use the trading pattern: WebSocket, in-memory state, movement detection.
Do you display data to a large audience? Use the media pattern: one fetcher, a cache and optionally widgets.
Are you training models or doing research? Use the analytics pattern: historical batch plus versioned datasets.
Building a fantasy product? Combine player statistics with a feature pipeline.</p>
<p>Most real products are hybrids. A betting media site, for example, uses the media pattern for scores and the sportsbook-style reference pattern for an odds board. That is fine. Combine patterns deliberately, and keep each in its own module.</p>
<ol>
<li>FAQ</li>
</ol>
<p>Is the API different for sportsbooks, trading desks and media sites?
No. The same normalized API serves every customer type. The Solutions pages highlight the parts most relevant to each use case, and the architecture around the API is what differs.</p>
<p>Which delivery method should I start with?
Start with REST to learn the data. Add WebSocket if you need continuous live streams, and webhooks for event-driven actions.</p>
<p>Can I build a media site without writing a front end?
Yes. Widgets are designed for that. Check the widgets page for what is currently available on your plan.</p>
<p>How do I avoid hitting rate limits on a high-traffic site?
Decouple reader traffic from API traffic with a cache, as in the media architecture. Your API usage then depends on how many matches you track, not how many people read.</p>
<p>Where do historical odds and closing odds fit?
They are for backtesting and research. See the Historical Sports Data API.</p>
<p>Where do I get started?
Create a free account, read the documentation, and try your first call in the sandbox.</p>
<p>Conclusion</p>
<p>The three Solutions pages are really three answers to one question: what is the hard problem in your system?</p>
<p>For a sportsbook, it is correctness: trustworthy reference prices and exactly-once settlement.
For a trading desk, it is speed and change detection: persistent streams, in-memory state and a clear backpressure policy.
For a media site, it is scale and resilience: fetch once, cache everywhere and keep readers away from your upstream.</p>
<p>Pick the pattern that matches your hard problem, keep sport-specific rules in an adapter layer, and let the normalized API do the heavy lifting underneath.</p>
<p>Ready to build? Get a free API key, explore the Odds API, Live Scores API and WebSocket API, or start from the Orbistats homepage.</p>
<p>Which architecture are you building, and where did it bite you? Tell me in the comments.</p>
]]></content:encoded></item><item><title><![CDATA[Orbistats API Versioning Explained: What /v1/ to /v2/ Actually Means for Your Integration]]></title><description><![CDATA[Every developer has lived this story. An integration works perfectly for months. Then one morning a field is renamed, a number turns into a string, or an enum gains a value your switch statement never]]></description><link>https://bettechmagnetics.hashnode.dev/orbistats-api-versioning-explained-what-v1-to-v2-actually-means-for-your-integration</link><guid isPermaLink="true">https://bettechmagnetics.hashnode.dev/orbistats-api-versioning-explained-what-v1-to-v2-actually-means-for-your-integration</guid><category><![CDATA[api versioning]]></category><category><![CDATA[REST API]]></category><category><![CDATA[API Design]]></category><category><![CDATA[backend]]></category><category><![CDATA[sports data]]></category><dc:creator><![CDATA[Vijay Choudhary]]></dc:creator><pubDate>Mon, 05 Oct 2026 13:30:06 GMT</pubDate><enclosure url="https://cdn.hashnode.com/uploads/covers/6a47ac8dec0b8a17b4dc915d/4dade707-48e0-44bf-b09d-cfdea8b5233d.png" length="0" type="image/jpeg"/><content:encoded><![CDATA[<p>Every developer has lived this story. An integration works perfectly for months. Then one morning a field is renamed, a number turns into a string, or an enum gains a value your switch statement never expected. Nothing in your code changed, but production is on fire.</p>
<p>API versioning exists to prevent that story. It is a contract between the people who run an API and the people who depend on it, and a contract is only useful if both sides understand it.</p>
<p>This post explains how versioning works on the Orbistats sports data and odds API: what /v1/ guarantees, what /v2/ would mean, and how to build an integration that treats upgrades as routine maintenance instead of an emergency.</p>
<p>Note: The code here is illustrative. For exact endpoints, fields and policies, always use the official documentation as your source of truth.</p>
<p>TL;DR
Orbistats exposes its current stable API under /v1/. Breaking changes ship under a new major version such as /v2/, not silently inside /v1/.
Non-breaking changes, such as new fields, can be added to the current version. Your code must tolerate extra data.
The version lives in the URL path, so it is explicit, cacheable and easy to see in logs.
The safest integration is a tolerant reader: it ignores unknown fields, handles unknown enum values and isolates the API behind a thin adapter layer.
Migration works best as a staged rollout: read the changelog, test in the sandbox, shadow-compare, canary, then cut over.
Table of Contents
Why API versioning matters
How Orbistats versions its API
Breaking vs non-breaking changes
What a /v1/ to /v2/ move might look like
Writing a tolerant reader
Isolating the API behind an adapter layer
Versioning across REST, WebSocket and webhooks
Versioning and sports data: why 13 sports make it harder
A safe migration playbook
Contract tests and CI checks
Shadow traffic: comparing v1 and v2 side by side
Monitoring: changelog, status and alerting
Common mistakes
FAQ</p>
<ol>
<li>Why API versioning matters</li>
</ol>
<p>A sports data API has an unusual property: your users see your mistakes live. If a score widget breaks during a match, or an odds board shows stale prices because parsing failed, it happens in front of an audience, in real time.</p>
<p>That raises the cost of every unexpected change. Versioning is how an API provider says: "We will keep this exact shape working, and when we need to change it, you will get advance notice and a separate version."</p>
<p>For you as the consumer, versioning gives you three things:</p>
<p>Predictability. The response you coded against keeps its shape.
Control over timing. You upgrade when you are ready, not when the provider is.
A clear signal. A new major version means "read the migration notes."
2. How Orbistats versions its API</p>
<p>According to the documentation, the current stable API lives under:</p>
<p>text
<a href="https://api.orbistats.com/v1/">https://api.orbistats.com/v1/</a></p>
<p>Authentication stays the same across the lifecycle of an integration:</p>
<p>http
GET <a href="https://api.orbistats.com/v1/fixtures">https://api.orbistats.com/v1/fixtures</a>
Authorization: Bearer YOUR_API_KEY</p>
<p>The model is simple:</p>
<p>Rule	What it means for you
Current stable version is /v1/	Build against this today
Breaking changes go to a new version, such as /v2/	Your /v1/ integration is not broken by them
Non-breaking fields may be added to the current version	Your parser must tolerate new fields</p>
<p>The version sits in the URL path rather than a header or query parameter. That choice has real advantages:</p>
<p>Visible. You can see the version in logs, browser dev tools, curl output and proxy dashboards.
Cache-friendly. CDNs and HTTP caches key on the URL, so /v1/ and /v2/ responses never collide.
Easy to route. Gateways, WAFs and API management tools can route on path prefixes with no extra logic.
Easy to test. Switching versions is a one-line change to a base URL.</p>
<p>The resources in the docs include fixtures, results, standings, odds, statistics, lineups, events, teams, players, competitions and countries. The full list is in the API Reference.</p>
<ol>
<li>Breaking vs non-breaking changes</li>
</ol>
<p>The entire versioning model rests on one distinction, so it is worth being precise about it.</p>
<p>Non-breaking (safe to add inside /v1/)
Adding a new field to a response object
Adding a new endpoint
Adding a new optional query parameter
Adding a new sport, competition or market to existing coverage
Improving performance, latency or data quality without changing the schema
Breaking (requires a new major version)
Removing or renaming a field
Changing a field's type, such as int to string, or an object to an array
Changing the meaning of a field (for example, minutes from "elapsed" to "remaining")
Making an optional parameter required
Changing authentication or error format in an incompatible way
Changing default behavior that existing clients rely on, such as default sort order or pagination size</p>
<p>A useful rule: if a correctly written client could stop working after the change, it is breaking. Extra fields do not break a correct client. A removed field does.</p>
<p>Notice the phrase "correctly written". Versioning is a two-sided deal. The provider promises not to break you, and you promise to write a client that tolerates additions. If your code crashes on an unknown field, no version policy can fully protect you.</p>
<ol>
<li>What a /v1/ to /v2/ move might look like</li>
</ol>
<p>Concrete examples make this tangible. Suppose /v1/ returns a live match like this:</p>
<p>json
{
  "match_id": "match_50231",
  "status": "live",
  "minute": 72,
  "home": { "name": "Manchester City", "score": 2 },
  "away": { "name": "Arsenal", "score": 1 }
}</p>
<p>Imagine a future /v2/ that needs to support more sports and richer clock data. The response might evolve like this:</p>
<p>json
{
  "match_id": "match_50231",
  "status": "live",
  "clock": { "elapsed": 72, "period": "2H", "added_time": 0 },
  "participants": [
    { "role": "home", "name": "Manchester City", "score": 2 },
    { "role": "away", "name": "Arsenal", "score": 1 }
  ]
}</p>
<p>Each difference is a classic breaking change:</p>
<p>Change	Why it breaks naive clients
minute becomes clock.elapsed	Field moved, so data["minute"] raises KeyError
home and away become a participants array	Object access turns into index access
period is introduced	Not breaking by itself, but a clock model now exists that old code ignores</p>
<p>This is hypothetical. It is not a statement about what /v2/ will contain. The point is the shape of the problem: structural improvements that are good for the platform are exactly the changes that cannot happen silently.</p>
<ol>
<li>Writing a tolerant reader</li>
</ol>
<p>The single most valuable habit for surviving any API upgrade is the tolerant reader pattern:</p>
<p>Be conservative in what you require, and liberal in what you accept.</p>
<p>In practice:</p>
<p>Read only the fields you need.
Ignore unknown fields.
Handle unknown enum values gracefully.
Treat optional data as optional.
Validate at the boundary and fail loudly there, not deep inside business logic.
Python example with explicit parsing
python
from dataclasses import dataclass
from typing import Optional</p>
<p>KNOWN_STATUSES = {"scheduled", "live", "finished", "postponed", "cancelled"}</p>
<p>@dataclass(frozen=True)
class Match:
    match_id: str
    status: str
    minute: Optional[int]
    home_name: str
    home_score: Optional[int]
    away_name: str
    away_score: Optional[int]</p>
<p>def parse_match_v1(raw: dict) -&gt; Match:
    status = raw.get("status", "unknown")
    if status not in KNOWN_STATUSES:
        # Unknown enum value: do NOT crash. Map to a safe default and log it.
        log_unknown_enum("match.status", status)
        status = "unknown"</p>
<pre><code>home = raw.get("home") or {}
away = raw.get("away") or {}

return Match(
    match_id=raw["match_id"],            # truly required: fail loudly if missing
    status=status,
    minute=raw.get("minute"),            # optional
    home_name=home.get("name", ""),
    home_score=home.get("score"),
    away_name=away.get("name", ""),
    away_score=away.get("score"),
)
</code></pre>
<p>Two details matter here. Extra fields in raw are simply never touched, so a new field added in /v1/ cannot hurt you. And match_id is required while everything else is optional, which keeps your failure modes intentional.</p>
<p>Strict-by-default parsers are a trap</p>
<p>Libraries that validate JSON into typed models can be configured two ways. In strict mode, unknown fields raise an error. In ignore mode, unknown fields are dropped silently. For third-party APIs you want ignore mode:</p>
<p>python
from pydantic import BaseModel, ConfigDict</p>
<p>class TeamSide(BaseModel):
    model_config = ConfigDict(extra="ignore")   # tolerate new fields
    name: str
    score: int | None = None</p>
<p>class MatchModel(BaseModel):
    model_config = ConfigDict(extra="ignore")
    match_id: str
    status: str
    minute: int | None = None
    home: TeamSide
    away: TeamSide</p>
<p>If you set extra="forbid" on an external API model, you have built a system that breaks the first time the provider adds a harmless field.</p>
<p>JavaScript / TypeScript equivalent
typescript
type MatchStatus = "scheduled" | "live" | "finished" | "postponed" | "cancelled" | "unknown";</p>
<p>const KNOWN: ReadonlySet = new Set([
  "scheduled", "live", "finished", "postponed", "cancelled",
]);</p>
<p>export function normalizeStatus(raw: unknown): MatchStatus {
  return typeof raw === "string" &amp;&amp; KNOWN.has(raw) ? (raw as MatchStatus) : "unknown";
}</p>
<p>export function parseMatchV1(raw: any) {
  if (!raw?.match_id) throw new Error("match_id missing"); // required field
  return {
    matchId: String(raw.match_id),
    status: normalizeStatus(raw.status),
    minute: typeof raw.minute === "number" ? raw.minute : null,
    home: { name: raw.home?.name ?? "", score: raw.home?.score ?? null },
    away: { name: raw.away?.name ?? "", score: raw.away?.score ?? null },
  };
}
6. Isolating the API behind an adapter layer</p>
<p>A tolerant reader protects you from additive changes. An adapter layer protects you from breaking ones, because it limits the blast radius of any migration to one folder.</p>
<p>The rule: only one module in your codebase knows what the Orbistats response looks like. Everything else uses your own domain model.</p>
<p>text
┌───────────────────────────┐
│  Your app / business code │   uses Match, Odds, Fixture (your types)
└─────────────┬─────────────┘
              │
┌─────────────▼─────────────┐
│    Domain model (yours)   │   stable, never changes with the API
└─────────────┬─────────────┘
              │
┌─────────────▼─────────────┐
│     Adapter layer         │   parse_match_v1 / parse_match_v2
└─────────────┬─────────────┘
              │
┌─────────────▼─────────────┐
│  HTTP client (versioned)  │   BASE = <a href="https://api.orbistats.com/v1">https://api.orbistats.com/v1</a>
└───────────────────────────┘</p>
<p>When /v2/ arrives, you write a second adapter and flip a switch. Nothing above the adapter changes.</p>
<p>python
import os
import requests</p>
<p>API_VERSION = os.getenv("ORBISTATS_API_VERSION", "v1")
BASE_URL = f"<a href="https://api.orbistats.com/%7BAPI_VERSION%7D">https://api.orbistats.com/{API_VERSION}</a>"</p>
<p>class OrbistatsClient:
    def <strong>init</strong>(self, api_key: str, session: requests.Session | None = None):
        self.session = session or requests.Session()
        self.session.headers.update({"Authorization": f"Bearer {api_key}"})</p>
<pre><code>def get(self, path: str, **params):
    resp = self.session.get(f"{BASE_URL}{path}", params=params, timeout=10)
    resp.raise_for_status()
    return resp.json()
</code></pre>
<h1>Parsers are registered per version</h1>
<p>PARSERS = {
    "v1": parse_match_v1,
    # "v2": parse_match_v2,   # added when you migrate
}</p>
<p>def get_live_match(client: OrbistatsClient, sport: str, match_id: str) -&gt; Match:
    raw = client.get(f"/{sport}/matches/{match_id}")
    return PARSERS<a href="raw">API_VERSION</a></p>
<p>With this structure, the version is configuration, not a code edit. You can run v1 in production and v2 in staging with the same codebase.</p>
<ol>
<li>Versioning across REST, WebSocket and webhooks</li>
</ol>
<p>REST is the easy case. Orbistats also delivers data through WebSocket and Webhooks, and these have different versioning concerns.</p>
<p>REST</p>
<p>You choose the version per request through the URL. Control is fine-grained.</p>
<p>WebSocket</p>
<p>A persistent connection is long-lived, so the schema of messages is fixed for the life of that connection. When you migrate, treat it like a deploy:</p>
<p>Open a second connection on the new version.
Run both in parallel and compare.
Drain and close the old connection.</p>
<p>Always implement reconnect with backoff, because a reconnect is also the natural moment to pick up a new version.</p>
<p>Webhooks</p>
<p>Webhooks are the most dangerous place for schema drift, because the provider is calling you, and a failing receiver means lost events. Build a receiver that can handle more than one payload shape:</p>
<p>python
from fastapi import FastAPI, Request</p>
<p>app = FastAPI()</p>
<p>HANDLERS = {
    1: handle_event_v1,
    # 2: handle_event_v2,
}</p>
<p>@app.post("/webhook/sports")
async def sports_hook(request: Request):
    payload = await request.json()</p>
<pre><code># Prefer an explicit schema/version marker if the payload provides one;
# otherwise fall back to the version you configured for your subscription.
version = payload.get("schema_version", 1)

handler = HANDLERS.get(version)
if handler is None:
    # Never 5xx on an unknown version: you'd trigger retries and lose visibility.
    log_unsupported_version(version, payload)
    return {"ok": True, "ignored": True}

handler(payload)
return {"ok": True}
</code></pre>
<p>Check the webhook docs for the exact version marker. The principle holds either way: accept, log and defer rather than crash.</p>
<ol>
<li>Versioning and sports data: why 13 sports make it harder</li>
</ol>
<p>Orbistats covers 13 sports: Football, Basketball, American Football, Cricket, Tennis, Baseball, Esports, Combat Sports, Volleyball, Handball, Ice Hockey, Golf and Horse Racing.</p>
<p>Why does that matter for versioning? Because a unified schema across very different sports is exactly the kind of design that eventually forces a major version.</p>
<p>Consider how differently sports model "the clock" and "the score":</p>
<p>Sport	Time model	Score model
Football	Two halves plus stoppage time	Goals
Basketball	Four quarters, countdown clock	Points
American Football	Four quarters, downs and possession	Points
Cricket	Overs and innings, no clock	Runs / wickets
Tennis	No clock: points, games, sets	Nested set scores
Baseball	Innings (top and bottom)	Runs
Esports	Maps / games in a series	Map wins
Combat Sports	Rounds	Result by method
Volleyball	Sets	Set and point scores
Handball	Two halves	Goals
Ice Hockey	Three periods plus OT	Goals
Golf	Rounds and holes	Strokes vs par
Horse Racing	Race, no score	Finishing position</p>
<p>A single field called minute makes sense for football and means nothing for tennis. As coverage grows, a platform will often need a richer, more general model, and that kind of redesign is a textbook reason to introduce /v2/ instead of bending /v1/ out of shape.</p>
<p>What this means for your integration: if you consume more than one sport, do not hard-code football assumptions into your domain model. Keep sport-specific parsing inside the adapter, so a future schema change in one sport does not ripple through your whole app.</p>
<ol>
<li>A safe migration playbook</li>
</ol>
<p>When a new major version is announced, resist the urge to either panic or ignore it. Follow a staged plan.</p>
<p>Stage 0: Read before you touch anything</p>
<p>Check the changelog and the migration notes in the documentation. Make a list of the endpoints you actually use. Most migrations affect only a fraction of any one customer's surface area.</p>
<p>Stage 1: Inventory your usage</p>
<p>Find every place your code calls the API. If you followed the adapter pattern, that is one directory. If not, this is a good moment to create the adapter.</p>
<p>bash</p>
<h1>Quick inventory of API call sites</h1>
<p>grep -rn "api.orbistats.com" --include="<em>.py" --include="</em>.ts" --include="*.js" .
Stage 2: Try the new version in the sandbox</p>
<p>Use the Sandbox to make requests against the new version and compare responses with /v1/ for the same resource. The Quickstart and Examples pages are good starting points.</p>
<p>Stage 3: Write the new adapter</p>
<p>Add parse_match_v2 (and friends) next to the v1 versions. Do not delete v1 yet.</p>
<p>Stage 4: Contract tests and shadow comparison</p>
<p>Run both versions side by side on real traffic. Sections 10 and 11 cover how.</p>
<p>Stage 5: Canary</p>
<p>Route a small percentage of traffic, or one non-critical feature, to the new version. Watch error rates, latency and parse failures.</p>
<p>Stage 6: Cut over, keep a rollback</p>
<p>Switch the version config. Keep the v1 adapter for at least one release cycle so you can roll back with a config change.</p>
<p>Stage 7: Clean up</p>
<p>Once you are comfortable and the old version is retired, delete the old adapter and tests.</p>
<p>If you use an official client, check the SDKs page for version support and pin your SDK version the same way you pin the API version.</p>
<ol>
<li>Contract tests and CI checks</li>
</ol>
<p>You cannot rely on remembering to check for changes. Automate it with contract tests: small tests that assert the parts of the response you depend on.</p>
<p>python</p>
<h1>tests/test_contract_orbistats.py</h1>
<p>import os
import pytest
import requests</p>
<p>BASE = f"<a href="https://api.orbistats.com/%7Bos.getenv">https://api.orbistats.com/{os.getenv</a>('ORBISTATS_API_VERSION', 'v1')}"
HEADERS = {"Authorization": f"Bearer {os.environ['ORBISTATS_API_KEY']}"}</p>
<p>REQUIRED_MATCH_FIELDS = {"match_id", "status"}</p>
<p>@pytest.mark.contract
def test_live_matches_contract():
    r = requests.get(f"{BASE}/football/matches/live", headers=HEADERS, timeout=10)
    assert r.status_code == 200</p>
<pre><code>body = r.json()
items = body["data"] if isinstance(body, dict) and "data" in body else body
assert isinstance(items, list)

for m in items[:20]:
    missing = REQUIRED_MATCH_FIELDS - set(m.keys())
    assert not missing, f"missing required fields: {missing}"
    assert isinstance(m["match_id"], str)
</code></pre>
<p>Run these tests:</p>
<p>On every deploy, to catch regressions you caused.
On a schedule (for example, nightly), to catch changes the provider caused.
Against the new version in staging before any migration.</p>
<p>A scheduled contract test is effectively an early-warning system. If the response shape changes, you find out from a failed test and not from a user report.</p>
<p>Because the free tier has request limits, keep contract tests small: a handful of calls, not hundreds. See pricing for current limits.</p>
<ol>
<li>Shadow traffic: comparing v1 and v2 side by side</li>
</ol>
<p>Contract tests prove the shape is right. Shadow comparison proves the meaning is right, which matters more.</p>
<p>The idea: for the same request, call both versions, normalize both through their adapters into your domain model, and diff the results.</p>
<p>python
import json
from deepdiff import DeepDiff</p>
<p>def shadow_compare(client_v1, client_v2, sport: str, match_id: str):
    raw_v1 = client_v1.get(f"/{sport}/matches/{match_id}")
    raw_v2 = client_v2.get(f"/{sport}/matches/{match_id}")</p>
<pre><code>m1 = parse_match_v1(raw_v1)
m2 = parse_match_v2(raw_v2)

# Compare normalized domain objects, not raw JSON
diff = DeepDiff(m1.__dict__, m2.__dict__, ignore_order=True)
if diff:
    print(f"[DIFF] {sport}/{match_id}: {json.dumps(diff, default=str, indent=2)}")
return diff
</code></pre>
<h1>Run across a sample of matches from different sports</h1>
<p>SAMPLE = [("football", "match_50231"), ("basketball", "match_77120")]
for sport, mid in SAMPLE:
    shadow_compare(client_v1, client_v2, sport, mid)</p>
<p>Comparing normalized domain objects rather than raw JSON is the key trick. The raw JSON will always differ between versions, because that is the point of a new version. What should not differ is the meaning: the same score, the same status, the same team.</p>
<p>If you find differences, decide for each whether it is:</p>
<p>Expected (a documented semantic change you need to adapt to),
A bug in your new adapter, or
Something to report to the provider.
12. Monitoring: changelog, status and alerting</p>
<p>Good versioning hygiene is partly about process. Three habits keep you ahead of surprises.</p>
<ol>
<li><p>Watch the changelog. The changelog is where changes are announced. Subscribe, or add a recurring calendar reminder, or have a small script check it weekly.</p>
</li>
<li><p>Watch the status page. When something looks wrong, the status page tells you quickly whether it is your code or an incident on the provider side.</p>
</li>
<li><p>Instrument your own parsing. Track metrics that reveal drift early:</p>
</li>
</ol>
<p>python
from collections import Counter</p>
<p>unknown_enum_counter = Counter()</p>
<p>def log_unknown_enum(field: str, value: str):
    unknown_enum_counter[(field, value)] += 1
    # emit to your metrics system: parse.unknown_enum{field=..., value=...}</p>
<p>def log_unsupported_version(version, payload):
    # emit to your metrics system: webhook.unsupported_version{version=...}
    pass</p>
<p>A spike in "unknown enum" or "missing field" events is often the first sign that the response shape changed. Alert on it. It costs almost nothing and can save hours of debugging.</p>
<p>Also log the API version on every request and in every error report. When you are debugging at 2 a.m., "which version was this?" should never be a question.</p>
<ol>
<li>Common mistakes
Strict parsing of external payloads. extra="forbid" or an exhaustive switch with no default case turns every additive change into an outage.
Hard-coding the version string in many places. One constant, driven by config, not fifty scattered URLs.
Letting API shapes leak into your whole codebase. If your UI reads response.home.score directly, a schema change touches every component.
Treating "new field" as dangerous. Additive changes are the good kind. Your code should shrug at them.
No contract tests. If the only thing that detects a change is a user complaint, you are testing in production.
Skipping the sandbox. Trying the new version against real code is cheap. Discovering a mismatch after cutover is not.
Migrating everything at once. Move one endpoint or feature at a time. Smaller steps mean smaller rollbacks.
Forgetting long-lived connections. WebSocket clients and webhook endpoints are easy to overlook during a migration.
Not pinning SDKs. An SDK auto-upgrading under you is versioning drift in disguise.
Ignoring semantic changes. A field with the same name but a different meaning is the nastiest kind of breaking change, and only shadow comparison reliably catches it.</li>
<li>FAQ</li>
</ol>
<p>What does /v1/ mean in the Orbistats API?
It is the current stable major version of the API. Your base URL is <a href="https://api.orbistats.com/v1/">https://api.orbistats.com/v1/</a>, and integrations built against it are protected from breaking changes, which are reserved for a new major version.</p>
<p>What happens when a /v2/ is released?
Breaking changes are introduced under the new version so existing /v1/ integrations are not silently broken. You migrate on your own schedule, guided by the changelog and documentation.</p>
<p>Can fields be added to /v1/ without a new version?
Yes. Non-breaking additions, such as new fields, can appear in the current version. That is why your client should ignore fields it does not recognize.</p>
<p>Do I need to change my API key or auth when I move versions?
Authentication uses a bearer token in the Authorization header. Check the docs for any version-specific notes before you migrate.</p>
<p>Why is the version in the URL and not in a header?
Path versioning is explicit and easy to cache, log and route. Everyone can see which version a request used just by reading the URL.</p>
<p>How much does migration cost me?
If you used a tolerant reader and an adapter layer, it is typically one new parser and a config change. If your code reads raw JSON everywhere, it is a larger refactor, which is a good argument for introducing the adapter layer now.</p>
<p>Where do I start?
Create a free key at signup, try requests in the Sandbox, and read the documentation.</p>
<p>Conclusion</p>
<p>Versioning is not just a URL prefix. It is a shared agreement: the provider keeps /v1/ stable and moves breaking changes into a new version, and you write a client that tolerates the safe changes and isolates the unsafe ones.</p>
<p>If you remember five things, make them these:</p>
<p>Pin the version in one place.
Parse tolerantly: ignore unknown fields, survive unknown enum values.
Put an adapter layer between the API and your business logic.
Back it with contract tests and shadow comparison.
Migrate in stages with a rollback ready.</p>
<p>Do that and a /v1/ to /v2/ move becomes a planned, boring, low-risk release, which is exactly what a good migration should feel like.</p>
<p>Ready to build? Get a free API key, explore the Odds API, Live Scores API and Sports Data API, or start from the Orbistats homepage.</p>
<p>Have a versioning horror story or a migration trick that saved you? Share it in the comments.</p>
]]></content:encoded></item><item><title><![CDATA[How Orbistats Normalizes Odds Across Bookmakers: A Technical Breakdown]]></title><description><![CDATA[If you have ever integrated more than one bookmaker feed, you already know the problem. Bookmaker A sends "1.91". Bookmaker B sends "-110". Bookmaker C sends "10/11". All three mean roughly the same t]]></description><link>https://bettechmagnetics.hashnode.dev/how-orbistats-normalizes-odds-across-bookmakers-a-technical-breakdown</link><guid isPermaLink="true">https://bettechmagnetics.hashnode.dev/how-orbistats-normalizes-odds-across-bookmakers-a-technical-breakdown</guid><category><![CDATA[Odds API]]></category><category><![CDATA[API Design]]></category><category><![CDATA[data-engineering]]></category><category><![CDATA[sports data]]></category><category><![CDATA[backend]]></category><dc:creator><![CDATA[Vijay Choudhary]]></dc:creator><pubDate>Mon, 05 Oct 2026 13:25:43 GMT</pubDate><enclosure url="https://cdn.hashnode.com/uploads/covers/6a47ac8dec0b8a17b4dc915d/5ac48451-cb83-4cfa-8d9a-d553cf914c69.png" length="0" type="image/jpeg"/><content:encoded><![CDATA[<p>If you have ever integrated more than one bookmaker feed, you already know the problem. Bookmaker A sends "1.91". Bookmaker B sends "-110". Bookmaker C sends "10/11". All three mean roughly the same thing. Your code still has to figure that out, for every market, in every sport, on every update.</p>
<p>Multiply that by dozens of bookmakers, hundreds of markets and 13 sports, and you are no longer building a product. You are maintaining parsers.</p>
<p>This post explains how odds normalization works under the hood and how we approach it at Orbistats. Each stage has code you can adapt to your own pipeline. If you would rather skip the plumbing, the Odds API delivers the end result.</p>
<p>Note: The code in this article is a simplified reference implementation. It illustrates the concepts and is not our production source.</p>
<p>TL;DR
Raw bookmaker data differs in price format, market naming, line conventions, team naming, timestamps and update behavior.
Normalization is a multi-stage pipeline: ingest, entity resolution, market mapping, price normalization, line normalization, validation, then delivery.
The output is one canonical schema you can consume through REST, WebSocket or webhooks, with no per-bookmaker parsers on your side.
The hardest parts are not the maths. They are event matching and market semantics.
Table of Contents
Why odds are hard to normalize
The high-level architecture
Stage 1: Ingestion adapters
Stage 2: Event entity resolution
Stage 3: Market taxonomy mapping
Stage 4: Price normalization
Stage 5: Line normalization (spreads, totals, Asian handicaps)
Stage 6: Validation, staleness and de-duplication
The canonical schema
Delivery: REST, WebSocket and webhooks
Sport-by-sport edge cases (all 13 sports)
Consuming normalized odds in 5 minutes
Common pitfalls
What you can build on top
FAQ</p>
<ol>
<li>Why odds are hard to normalize</li>
</ol>
<p>"Just convert them to decimal" is the first idea most teams have. It covers about 10% of the problem. Here is the full list of differences.</p>
<p>Dimension	Example of variation
Price format	Decimal 1.91, American -110, fractional 10/11, Hong Kong 0.91, Indonesian, Malay
Market naming	1X2, Match Result, Full Time Result, Moneyline, Home/Draw/Away
Outcome naming	Home/1/team name, Over/O/Total Over 2.5
Line conventions	Spread -1.5 on home vs +1.5 on away, quarter lines such as -0.25
Participant naming	Man City, Manchester City FC, M. City
Event identity	Each bookmaker has its own event ID, and kickoff times can differ by minutes
Timestamps	Server time, local time, no timestamp at all
Update behavior	Full snapshots vs deltas, suspended markets, silent removals</p>
<p>Each of these is a failure mode. A normalizer that only handles price formats will still show your users a "best price" from a market that is actually a different bet.</p>
<ol>
<li>The high-level architecture
text
 Bookmaker feeds (heterogeneous)
 ┌────────┐ ┌────────┐ ┌────────┐ ┌────────┐
 │ Book A │ │ Book B │ │ Book C │ │ Book N │
 └───┬────┘ └───┬────┘ └───┬────┘ └───┬────┘
  │          │          │          │
  ▼          ▼          ▼          ▼
 ┌──────────────────────────────────────────┐
 │  1. Ingestion adapters (per source)      │
 └───────────────────┬──────────────────────┘
              ▼</li>
</ol>
<p> ┌──────────────────────────────────────────┐
 │  2. Event entity resolution              │
 │     (which events are the same match?)   │
 └───────────────────┬──────────────────────┘
                     ▼
 ┌──────────────────────────────────────────┐
 │  3. Market taxonomy mapping              │
 │     (which markets are the same bet?)    │
 └───────────────────┬──────────────────────┘
                     ▼
 ┌──────────────────────────────────────────┐
 │  4. Price normalization (decimal, prob.) │
 │  5. Line normalization (spreads/totals)  │
 └───────────────────┬──────────────────────┘
                     ▼
 ┌──────────────────────────────────────────┐
 │  6. Validation, staleness, de-dup        │
 └───────────────────┬──────────────────────┘
                     ▼
        Canonical odds store + change log
                     │
      ┌──────────────┼──────────────┐
      ▼              ▼              ▼
     REST        WebSocket       Webhooks</p>
<p>The key idea is that adapters are the only place where bookmaker-specific code lives. Everything after stage 1 operates on a common internal model. Adding a new bookmaker means writing one adapter, not touching the rest of the system.</p>
<ol>
<li>Stage 1: Ingestion adapters</li>
</ol>
<p>An adapter's only job is to convert a source's raw payload into an intermediate raw record. It does not clean, convert or interpret anything yet.</p>
<p>python
from dataclasses import dataclass
from datetime import datetime
from typing import Optional</p>
<p>@dataclass(frozen=True)
class RawQuote:
    source: str                  # "book_a"
    source_event_id: str
    source_market: str           # raw market label, untouched
    source_outcome: str          # raw outcome label, untouched
    price_raw: str               # untouched, may be "1.91", "-110", "10/11"
    line_raw: Optional[str]      # "-1.5", "2.5", "-0.25", None
    home_raw: str
    away_raw: str
    start_time_raw: str
    sport_raw: str
    received_at: datetime        # when WE received it, in UTC
    source_ts: Optional[datetime]  # when the source says it was priced</p>
<p>Two design decisions matter here.</p>
<ol>
<li><p>Keep raw values untouched. If a normalization bug surfaces three weeks later, you can replay the raw stream and reprocess it. If you only stored cleaned data, the evidence is gone.</p>
</li>
<li><p>Always stamp received_at yourself. Sources lie about time, or omit it. Your own clock is the only one you control, and you will need it for staleness detection later.</p>
</li>
<li><p>Stage 2: Event entity resolution</p>
</li>
</ol>
<p>This is the hardest stage. Before you can compare prices, you need to know that "Man City vs Arsenal" from Book A and "Manchester City FC v Arsenal FC" from Book B are the same match.</p>
<p>A reliable approach combines three signals:</p>
<p>Sport and competition (hard filter)
Kickoff time within a tolerance window
Participant similarity after alias normalization
python
import re
import unicodedata
from difflib import SequenceMatcher</p>
<p>STOPWORDS = {"fc", "cf", "sc", "afc", "the", "club", "de", "ac"}</p>
<p>def canon_name(name: str) -&gt; str:
    name = unicodedata.normalize("NFKD", name).encode("ascii", "ignore").decode()
    name = name.lower()
    name = re.sub(r"[^a-z0-9 ]", " ", name)
    tokens = [t for t in name.split() if t not in STOPWORDS]
    return " ".join(tokens)</p>
<p>def name_score(a: str, b: str, aliases: dict) -&gt; float:
    ca, cb = canon_name(a), canon_name(b)
    # Explicit alias table wins over fuzzy matching
    if aliases.get(ca) == cb or aliases.get(cb) == ca:
        return 1.0
    return SequenceMatcher(None, ca, cb).ratio()</p>
<p>def same_event(ev_a, ev_b, aliases, max_minutes=15, threshold=0.82) -&gt; bool:
    if ev_a.sport != ev_b.sport:
        return False
    delta = abs((ev_a.start_utc - ev_b.start_utc).total_seconds()) / 60
    if delta &gt; max_minutes:
        return False
    home = name_score(ev_a.home, ev_b.home, aliases)
    away = name_score(ev_a.away, ev_b.away, aliases)
    # Also try swapped order: some sources flip home/away for neutral venues
    home_sw = name_score(ev_a.home, ev_b.away, aliases)
    away_sw = name_score(ev_a.away, ev_b.home, aliases)
    straight = min(home, away)
    swapped = min(home_sw, away_sw)
    return max(straight, swapped) &gt;= threshold</p>
<p>In production you usually go further:</p>
<p>Maintain a curated alias table that grows over time. Fuzzy matching alone produces false positives such as "Real Madrid" vs "Real Madrid Castilla".
Use a blocking step (group by sport, date and league) so you never compare every event with every other event.
Flag low-confidence matches for review instead of silently merging them. A wrong merge poisons prices for both events.
Once matched, assign a canonical event_id. All downstream data keys off this ID, never the source's ID.</p>
<p>A good rule is that a missed match costs you coverage, but a wrong match costs you correctness. Prefer the former.</p>
<ol>
<li>Stage 3: Market taxonomy mapping</li>
</ol>
<p>Next, you map every source's market label to a canonical market type. This is where semantics matter, because the same label can mean different bets.</p>
<p>For example, "Match Result" in football typically means 90 minutes plus stoppage time. In basketball, "Moneyline" includes overtime. A mapper that ignores that will compare unlike bets.</p>
<p>python
from enum import Enum</p>
<p>class Market(str, Enum):
    ML_3WAY = "1X2"           # home / draw / away, regulation time
    ML_2WAY = "MONEYLINE"     # home / away, incl. OT / extra time
    SPREAD = "SPREAD"         # handicap / point spread
    TOTAL = "TOTAL"           # over / under
    PLAYER_PROP = "PLAYER_PROP"
    FUTURE = "FUTURE"</p>
<h1>(sport, normalized source label) -&gt; canonical market + period + scope</h1>
<p>MARKET_MAP = {
    ("football", "match result"):        (Market.ML_3WAY, "regulation"),
    ("football", "full time result"):    (Market.ML_3WAY, "regulation"),
    ("football", "1x2"):                 (Market.ML_3WAY, "regulation"),
    ("basketball", "moneyline"):         (Market.ML_2WAY, "incl_overtime"),
    ("basketball", "money line"):        (Market.ML_2WAY, "incl_overtime"),
    ("tennis", "match winner"):          (Market.ML_2WAY, "match"),
    ("ice-hockey", "moneyline"):         (Market.ML_2WAY, "incl_overtime"),
    ("football", "asian handicap"):      (Market.SPREAD, "regulation"),
    ("basketball", "point spread"):      (Market.SPREAD, "incl_overtime"),
    ("football", "goals over/under"):    (Market.TOTAL, "regulation"),
}</p>
<p>def map_market(sport: str, raw_label: str):
    key = (sport, raw_label.strip().lower())
    try:
        return MARKET_MAP[key]
    except KeyError:
        # Unmapped markets are logged, never guessed
        raise UnmappedMarket(sport, raw_label)</p>
<p>Two rules keep this stage honest:</p>
<p>Unmapped means dropped and logged, not guessed. A silent wrong mapping is worse than a missing market.
Always carry the period/scope (regulation, incl. overtime, first half, and so on). Two markets with the same type but different periods are different bets.</p>
<p>Outcome labels get the same treatment. "1", "Home", "Manchester City" and "Team 1" all resolve to the canonical outcome home, and you determine that after event resolution, because the participant names are what let you assign home and away.</p>
<ol>
<li>Stage 4: Price normalization</li>
</ol>
<p>With events and markets resolved, price conversion is the easy part. Always convert to decimal odds as the canonical form, and always keep the implied probability next to it.</p>
<p>python
from fractions import Fraction</p>
<p>def to_decimal(price: str) -&gt; float:
    p = price.strip()</p>
<pre><code># Fractional: "10/11"
if "/" in p:
    frac = Fraction(p)
    return float(1 + frac)

val = float(p)

# American: values like -110, +150 (abs &gt;= 100)
if abs(val) &gt;= 100:
    if val &gt; 0:
        return 1 + val / 100
    return 1 + 100 / abs(val)

# Decimal (european): 1.01 .. 1000
if val &gt; 1.0:
    return val

# Hong Kong: stake not included, 0.91 -&gt; 1.91
return 1 + val
</code></pre>
<p>Caveat: Hong Kong, Malay and Indonesian formats are ambiguous when detected from the number alone. In production, take the format from the adapter's source configuration instead of inferring it from the value. The function above is intentionally simple.</p>
<p>Implied probability and margin (overround)</p>
<p>Every bookmaker price embeds a margin. For a market with outcomes i, the booksum is:</p>
<p>text
booksum = Σ (1 / decimal_odds_i)</p>
<p>If the booksum is above 1.0, the excess is the bookmaker's margin. Removing it gives you "fair" probabilities, which matter for modeling and for comparing bookmakers fairly.</p>
<p>python
def implied_probs(decimals: list[float]) -&gt; list[float]:
    return [1 / d for d in decimals]</p>
<p>def overround(decimals: list[float]) -&gt; float:
    return sum(implied_probs(decimals)) - 1.0</p>
<p>def devig_proportional(decimals: list[float]) -&gt; list[float]:
    """Simple proportional margin removal."""
    probs = implied_probs(decimals)
    total = sum(probs)
    return [p / total for p in probs]</p>
<h1>Example: 1X2 market</h1>
<p>odds = [1.91, 3.40, 4.20]
print(round(overround(odds) * 100, 2), "% margin")      # about 5.6 % margin
print([round(p, 4) for p in devig_proportional(odds)])</p>
<p>Proportional removal is the baseline. More advanced methods (power method, Shin's method) handle favorite-longshot bias better. Whichever you use, store both the raw implied probability and the devigged one so your users can choose.</p>
<ol>
<li>Stage 5: Line normalization (spreads, totals, Asian handicaps)</li>
</ol>
<p>Lines have two classic traps: sign conventions and quarter lines.</p>
<p>Sign convention</p>
<p>Pick one convention and enforce it. A good canonical rule is that the line is always expressed from the home team's perspective, and the away outcome gets the inverse.</p>
<p>python
def normalize_spread(home_line: float):
    """Home -1.5 is equivalent to Away +1.5."""
    return {"home": home_line, "away": -home_line}
Quarter lines (Asian handicap)</p>
<p>A line such as -0.25 is a split bet: half the stake on 0 and half on -0.5. If you do not decompose it, you cannot compare it with a bookmaker that offers only half-lines.</p>
<p>python
def split_quarter_line(line: float):
    """-0.25 -&gt; [(-0.0, 0.5), (-0.5, 0.5)]; -0.5 -&gt; [(-0.5, 1.0)]"""
    frac = abs(line) % 1
    sign = -1 if line &lt; 0 else 1
    base = int(abs(line))</p>
<pre><code>if frac == 0.25:
    return [(sign * base, 0.5), (sign * (base + 0.5), 0.5)]
if frac == 0.75:
    return [(sign * (base + 0.5), 0.5), (sign * (base + 1), 0.5)]
return [(line, 1.0)]
</code></pre>
<p>print(split_quarter_line(-0.25))   # [(-0, 0.5), (-0.5, 0.5)]
print(split_quarter_line(-0.75))   # [(-0.5, 0.5), (-1, 0.5)]</p>
<p>Totals follow the same idea. Over 2.5 and Over 2.25 are different products, so your canonical key must include the line value.</p>
<ol>
<li>Stage 6: Validation, staleness and de-duplication</li>
</ol>
<p>Normalized data is only useful if it is trustworthy. A validation layer sits between the mapping stages and the store.</p>
<p>python
from datetime import datetime, timedelta, timezone</p>
<p>MAX_AGE_PREMATCH = timedelta(seconds=120)
MAX_AGE_LIVE = timedelta(seconds=5)</p>
<p>def validate_quote(q, now=None) -&gt; list[str]:
    now = now or datetime.now(timezone.utc)
    problems = []</p>
<pre><code># 1. Price sanity
if not (1.01 &lt;= q.decimal &lt;= 1000):
    problems.append("price_out_of_range")

# 2. Staleness
max_age = MAX_AGE_LIVE if q.is_live else MAX_AGE_PREMATCH
if now - q.received_at &gt; max_age:
    problems.append("stale")

# 3. Market-level sanity (booksum shouldn't be absurd)
if q.market_booksum is not None and not (0.95 &lt;= q.market_booksum &lt;= 1.35):
    problems.append("booksum_outlier")

return problems
</code></pre>
<p>Other checks worth having:</p>
<p>Suspension handling. If a source suspends a market (goal scored, red card, injury), propagate a suspended state instead of leaving the last price on screen.
Idempotent updates. Store a hash of (event, market, line, outcome, bookmaker, price). If the hash did not change, do not emit an update.
Outlier detection. A price 40% off the market consensus is usually a data error, but it can also be a real stale-line opportunity. Flag it, do not delete it.
Monotonic ordering per key. Ignore updates older than the one you already hold.
9. The canonical schema</p>
<p>After all stages, every quote lands in one shape. This is what makes the downstream API predictable.</p>
<p>json
{
  "event_id": "evt_8f2c91",
  "sport": "football",
  "competition": "premier-league",
  "start_time": "2026-10-15T19:00:00Z",
  "home": { "id": "team_mci", "name": "Manchester City" },
  "away": { "id": "team_ars", "name": "Arsenal" },
  "market": "1X2",
  "period": "regulation",
  "line": null,
  "is_live": false,
  "bookmakers": [
    {
      "bookmaker": "book_a",
      "updated_at": "2026-10-15T14:02:11.482Z",
      "status": "open",
      "outcomes": {
        "home": { "decimal": 1.91, "implied_prob": 0.5236, "fair_prob": 0.4955 },
        "draw": { "decimal": 3.40, "implied_prob": 0.2941, "fair_prob": 0.2783 },
        "away": { "decimal": 4.20, "implied_prob": 0.2381, "fair_prob": 0.2253 }
      },
      "overround": 0.0558
    }
  ]
}</p>
<p>Notice what is not in there: nothing bookmaker-specific leaks into the structure. The same shape serves a football 1X2, a basketball moneyline or a tennis match winner. Only the market, period and line fields change.</p>
<ol>
<li>Delivery: REST, WebSocket and webhooks</li>
</ol>
<p>A normalized dataset has to reach your application in the way that fits your workload. Orbistats exposes it three ways.</p>
<p>REST for snapshots and backfills
http
GET <a href="https://api.orbistats.com/v1/odds?sport=football&amp;market=1X2">https://api.orbistats.com/v1/odds?sport=football&amp;market=1X2</a>
Authorization: Bearer YOUR_API_KEY</p>
<p>Use REST for page loads, dashboards, scheduled jobs and analysis. The API Reference lists every parameter.</p>
<p>WebSocket for live odds</p>
<p>Polling a live market every second wastes requests and still lags. With a persistent connection, the server pushes each change the moment it is normalized. Read more on the WebSocket API page.</p>
<p>javascript
// Illustrative client. Check the docs for the exact endpoint and message format.
const ws = new WebSocket("wss://api.orbistats.com/v1/stream?token=YOUR_API_KEY");</p>
<p>ws.onopen = () =&gt; {
  ws.send(JSON.stringify({
    action: "subscribe",
    channel: "odds",
    sport: "football",
    market: "1X2"
  }));
};</p>
<p>ws.onmessage = (msg) =&gt; {
  const update = JSON.parse(msg.data);
  console.log(update.event_id, update.bookmaker, update.outcomes);
};</p>
<p>ws.onclose = () =&gt; {
  // always reconnect with exponential backoff + resubscribe
};
Webhooks for event-driven workflows</p>
<p>If you do not want a long-lived connection, Webhooks push events to your server:</p>
<p>text
Odds move &gt; threshold
        ↓
Orbistats
        ↓
POST <a href="https://your-app.com/webhook/odds">https://your-app.com/webhook/odds</a>
        ↓
Your server (alert, trade, update UI)</p>
<p>A minimal receiver:</p>
<p>python
from fastapi import FastAPI, Request, HTTPException</p>
<p>app = FastAPI()</p>
<p>@app.post("/webhook/odds")
async def odds_hook(request: Request):
    payload = await request.json()
    # verify the signature header per the webhook docs before trusting the body
    if payload.get("type") == "odds.moved":
        handle_move(payload["event_id"], payload["market"], payload["delta"])
    return {"ok": True}</p>
<p>Return a 2xx quickly, push heavy work onto a queue, and make your handler idempotent, because any at-least-once delivery system can send the same event twice.</p>
<ol>
<li>Sport-by-sport edge cases (all 13 sports)</li>
</ol>
<p>Normalization rules are not sport-agnostic. Here is where each sport bites. Orbistats currently covers 13 sports.</p>
<p>Sport	Normalization gotchas
Football	3-way markets, regulation vs extra time, quarter-line Asian handicaps
Basketball	Moneyline includes overtime, large totals lines, many alternate spreads
American Football	Key numbers (3, 7) in spreads, moneyline includes OT, bye weeks
Cricket	Match format (T20, ODI, Test), draw possibility in Tests, rain interruptions
Tennis	Retirement and walkover rules differ per bookmaker, set and game handicaps
Baseball	Pitcher-dependent markets ("listed pitchers" action), run lines at ±1.5
Esports	Best-of-N series vs single map, game-specific markets
Combat Sports	Method-of-victory and round markets, draw and no-contest rules
Volleyball	Set handicaps, points totals per set
Handball	Draw handling, high-scoring totals
Ice Hockey	Regulation vs incl. OT moneyline, puck lines
Golf	Outright and each-way markets, dead-heat rules
Horse Racing	Each-way terms, non-runners, Starting Price vs fixed odds</p>
<p>The pattern across all of them is the same. The price is rarely the problem. The settlement rules are. Two bookmakers can both offer "Match Winner" for tennis and still treat a retirement differently. A good normalizer encodes that difference as metadata (rules.retirement = "void") rather than hiding it.</p>
<ol>
<li>Consuming normalized odds in 5 minutes</li>
</ol>
<p>Here is the end-to-end flow from the consumer side. No bookmaker parsers, no format conversion.</p>
<p>Step 1. Create a free account and get your API key.</p>
<p>Step 2. Try a request in the Sandbox or follow the Quickstart.</p>
<p>Step 3. Call it from your code:</p>
<p>python
import requests</p>
<p>API_KEY = "YOUR_API_KEY"
BASE = "<a href="https://api.orbistats.com/v1">https://api.orbistats.com/v1</a>"</p>
<p>resp = requests.get(
    f"{BASE}/odds",
    headers={"Authorization": f"Bearer {API_KEY}"},
    params={"sport": "football", "market": "1X2"},
    timeout=10,
)
resp.raise_for_status()</p>
<p>for event in resp.json().get("data", []):
    best = {}
    for book in event["bookmakers"]:
        for outcome, o in book["outcomes"].items():
            if outcome not in best or o["decimal"] &gt; best[outcome]["decimal"]:
                best[outcome] = {"bookmaker": book["bookmaker"], "decimal": o["decimal"]}
    print(event["home"]["name"], "vs", event["away"]["name"], best)</p>
<p>That 10-line "best price" loop is only possible because the data is normalized. Without normalization, it would be hundreds of lines of per-bookmaker handling.</p>
<p>There are also SDKs and code examples if you prefer a client library.</p>
<ol>
<li>Common pitfalls (and how to avoid them)
Comparing unlike markets. Always check market, period and line before comparing prices.
Ignoring staleness. A "best price" that is 90 seconds old on a live market is not a price. It is history. Use updated_at.
Trusting a single source of truth for names. Keep aliases in your own data layer if you join with other datasets.
Polling live markets. It is slow and expensive. Use WebSocket or webhooks.
Not storing history. Without a change log you cannot compute line movement or backtest. Use the Historical Sports Data API for multi-season archives, including closing odds.
Forgetting margin. Raw implied probabilities sum to more than 100%. Devig before you model.
Hard-coding bookmaker names. Treat bookmakers as data, not as enum values in your code.</li>
<li>What you can build on top</li>
</ol>
<p>Once the data is normalized, entire product categories open up:</p>
<p>Odds comparison sites that show the best price per outcome across bookmakers.
Line movement trackers that alert on sharp moves, which is a natural fit for trading desks.
Sportsbook back-office tools that benchmark your prices against the market. See Sportsbooks &amp; Trading.
Backtesting and ML pipelines that train on years of closing odds. See Analytics &amp; Data Science.
Live match centers that combine odds with the Live Scores API, fixtures from the Sports Data API and team and player numbers from the Sports Statistics API.</p>
<p>If you are new to the vocabulary (implied probability, overround, closing line), the glossary and guides cover the fundamentals.</p>
<ol>
<li>FAQ</li>
</ol>
<p>What does "normalized odds" actually mean?
Odds from many bookmakers converted into one consistent structure, with one price format, one market taxonomy, one line convention and one set of identifiers, so they can be compared directly.</p>
<p>Why not just convert everything to decimal?
Decimal conversion handles only price format. You still need to match events, map markets and unify line conventions, which are the harder problems.</p>
<p>How do I get the lowest latency?
Use WebSocket for live markets. REST is better for snapshots, and webhooks suit event-driven workflows.</p>
<p>Can I use this for backtesting?
Yes. Historical data, including closing odds, is available through the historical API. Check coverage per sport on the relevant page.</p>
<p>Is there a free tier?
Yes. You can start on the free tier and move up as you scale. See pricing for current plans and limits.</p>
<p>Conclusion</p>
<p>Odds normalization looks like a formatting problem, but it is really a data modeling and entity resolution problem. The price conversion is maybe 10% of the work. The rest is deciding which events are the same, which markets are comparable, which lines are equivalent and which quotes are still alive.</p>
<p>Getting that right once, centrally, is what lets everyone downstream ship faster. You write product code instead of parsers.</p>
<p>Ready to try it? Grab a free API key, read the documentation, and make your first call to the Odds API. Explore everything else at Orbistats.</p>
<p>Have a normalization edge case that bit you? Drop it in the comments. I would love to hear what you have run into.</p>
]]></content:encoded></item><item><title><![CDATA[From API to App: How Sports Data Reaches Your Screen]]></title><description><![CDATA[The part tutorials skip
Most sports API tutorials end at the same place: you call an endpoint, you console.log the JSON, and you feel great. Then you try to put it in a real app and things fall apart.]]></description><link>https://bettechmagnetics.hashnode.dev/from-api-to-app-how-sports-data-reaches-your-screen</link><guid isPermaLink="true">https://bettechmagnetics.hashnode.dev/from-api-to-app-how-sports-data-reaches-your-screen</guid><category><![CDATA[sports tech]]></category><category><![CDATA[React]]></category><category><![CDATA[api]]></category><category><![CDATA[JavaScript]]></category><category><![CDATA[Web Development]]></category><dc:creator><![CDATA[Vijay Choudhary]]></dc:creator><pubDate>Sat, 03 Oct 2026 19:43:49 GMT</pubDate><enclosure url="https://cdn.hashnode.com/uploads/covers/6a47ac8dec0b8a17b4dc915d/8fbccfee-37bc-47d4-a2b5-d3c41ea00665.png" length="0" type="image/jpeg"/><content:encoded><![CDATA[<p>The part tutorials skip</p>
<p>Most sports API tutorials end at the same place: you call an endpoint, you console.log the JSON, and you feel great. Then you try to put it in a real app and things fall apart.</p>
<p>The API key is sitting in your browser code where anyone can copy it. The browser blocks the request with a CORS error. Your free quota disappears in an hour because ten users opened the page. The score flickers. A failed request leaves a blank screen. And the kickoff time shows 3:30 AM for someone in another country.</p>
<p>None of these are "API problems". They are last-mile problems: everything between the data provider and the pixels on a user's screen. This article walks through that last mile, step by step, with code you can run.</p>
<p>You will learn:</p>
<p>Why your app should never call the sports API directly from the browser
How to build a small proxy server that hides your key and caches responses
How to shape the data so the frontend stays simple
How to write a React hook for live updates that behaves well on bad networks
How to design loading, stale and error states
How to show times correctly for any user
How to upgrade from polling to streaming later
Security and performance habits worth building from day one</p>
<p>You need basic JavaScript. React knowledge helps but is not required.</p>
<p>The journey in one picture
Sports data provider (REST / streaming)
            |
            v
   Your backend proxy (Node.js)</p>
<ul>
<li>holds the API key</li>
<li>caches responses</li>
<li>trims and reshapes data
   |
   v</li>
</ul>
<p>   Your frontend (React)</p>
<ul>
<li>polls or listens for updates</li>
<li>handles loading / stale / error</li>
<li>formats times and scores
   |
   v
   The user's screen</li>
</ul>
<p>Notice the proxy in the middle. That small piece is the difference between a demo and an app.</p>
<p>Why you should never call the provider from the browser</p>
<p>It is tempting. fetch works in the browser, the API gives you JSON, why add a server? Four reasons:</p>
<ol>
<li><p>Your API key becomes public. Anything shipped to the browser can be read by anyone with developer tools. Someone copies your key and burns your quota, or your bill. The Twelve-Factor App guidance on config is a good short read on keeping secrets out of code.</p>
</li>
<li><p>CORS will block you. Browsers enforce Cross-Origin Resource Sharing rules. Many providers do not allow arbitrary websites to call them directly, so the request fails before it starts.</p>
</li>
<li><p>Your quota gets multiplied. If 1,000 users each poll the provider every 5 seconds, that is 200 requests per second against your limit. A proxy fetches once and serves everyone.</p>
</li>
<li><p>You lose control of the data shape. The provider's response might be huge, messy or change over time. A proxy lets you hand the frontend only what it needs.</p>
</li>
</ol>
<p>The standard name for this pattern is a backend that serves one specific frontend. You will see it called a BFF ("backend for frontend"), but it is really just a small server.</p>
<p>Step 1: A tiny proxy server in Node.js</p>
<p>You need Node.js 18 or newer (it includes fetch) and Express.</p>
<p>bash
mkdir live-board &amp;&amp; cd live-board
npm init -y
npm install express</p>
<p>In package.json add "type": "module", then create server.js:</p>
<p>javascript
import express from "express";</p>
<p>const app = express();
const API_KEY = process.env.SPORTS_API_KEY;
const UPSTREAM = "<a href="https://api.example.com/v1">https://api.example.com/v1</a>";      // replace with your provider
const ALLOWED_SPORTS = new Set(["football", "cricket", "basketball", "tennis"]);</p>
<p>const cache = new Map();   // path -&gt; { data, expires }</p>
<p>async function cachedFetch(path, ttlMs) {
  const hit = cache.get(path);
  if (hit &amp;&amp; hit.expires &gt; Date.now()) return hit.data;</p>
<p>  const res = await fetch(<code>${UPSTREAM}${path}</code>, {
    headers: { Authorization: <code>Bearer ${API_KEY}</code> },
    signal: AbortSignal.timeout(8000),               // never wait forever
  });
  if (!res.ok) throw new Error(<code>Upstream responded ${res.status}</code>);</p>
<p>  const data = await res.json();
  cache.set(path, { data, expires: Date.now() + ttlMs });
  return data;
}</p>
<p>// Keep only what the frontend needs
function toCard(m) {
  return {
    id: m.id,
    status: m.status,
    minute: m.minute ?? null,
    kickoff: m.starts_at,                             // ISO 8601, UTC
    home: { name: m.home.name, score: m.home.score },
    away: { name: m.away.name, score: m.away.score },
  };
}</p>
<p>app.get("/api/matches/live", async (req, res) =&gt; {
  const sport = String(req.query.sport || "football");
  if (!ALLOWED_SPORTS.has(sport)) {
    return res.status(400).json({ error: "Unknown sport" });
  }</p>
<p>  try {
    const raw = await cachedFetch(<code>/matches?sport=${sport}&amp;status=live</code>, 3000);
    res.set("Cache-Control", "public, max-age=3");
    res.json({ data: raw.data.map(toCard), fetchedAt: new Date().toISOString() });
  } catch (err) {
    console.error(err.message);
    res.status(502).json({ error: "Live data is temporarily unavailable" });
  }
});</p>
<p>app.listen(3000, () =&gt; console.log("Proxy running on <a href="http://localhost:3000">http://localhost:3000</a>"));</p>
<p>Run it with your key kept outside the code:</p>
<p>bash
SPORTS_API_KEY=your_key_here node server.js</p>
<p>Every line here earns its place:</p>
<p>ALLOWED_SPORTS is a whitelist. Never paste user input straight into an upstream URL. This is one of the basic habits from the OWASP API Security project.
A 3-second in-memory cache means a thousand users in the same moment cost you one upstream request.
AbortSignal.timeout stops a slow provider from freezing your server.
toCard shrinks the payload. Less data means faster loads and less to break.
A friendly 502 message means the browser never sees your provider's internal error details.
Cache-Control tells browsers and CDNs how long a response stays fresh. MDN explains the options in its Cache-Control reference.</p>
<p>For a real deployment you would swap the Map for a shared cache, but the idea is identical.</p>
<p>Step 2: Choosing how fresh "live" should be</p>
<p>A live score does not need the same freshness as a stock ticker. Use different cache times for different data:</p>
<p>Data	Suggested freshness
Teams, leagues, venues	Hours to days
Fixtures, standings	30 seconds to a few minutes
Live score and clock	2 to 5 seconds</p>
<p>Check what your provider allows before choosing. Documentation from vendors such as SportsDataIO and Sportradar spells out refresh expectations and limits, and it is worth reading before you design your polling interval. If you want a multi-sport option with one consistent response structure, Orbistats is built for that, with 13 sports covered ([add your 13 sports here]), which means the same proxy and frontend code can serve every sport.</p>
<p>Step 3: A React hook for live data</p>
<p>Now the frontend. We will write a reusable hook with React that polls the proxy, handles failures, and slows down when the network is struggling.</p>
<p>jsx
import { useEffect, useRef, useState } from "react";</p>
<p>export function useLiveMatches(sport, intervalMs = 5000) {
  const [matches, setMatches] = useState([]);
  const [status, setStatus] = useState("loading");   // loading | live | stale | error
  const [updatedAt, setUpdatedAt] = useState(null);
  const failures = useRef(0);</p>
<p>  useEffect(() =&gt; {
    let cancelled = false;
    let timer;
    const controller = new AbortController();</p>
<pre><code>function schedule() {
  if (cancelled) return;
  // after failures, wait longer: 5s, 10s, 20s ... capped at 60s
  const wait = Math.min(intervalMs * 2 ** failures.current, 60_000);
  timer = setTimeout(load, wait);
}

async function load() {
  if (document.hidden) return schedule();         // tab not visible: skip this round

  try {
    const res = await fetch(`/api/matches/live?sport=${sport}`, {
      signal: controller.signal,
    });
    if (!res.ok) throw new Error(`HTTP ${res.status}`);
    const body = await res.json();
    if (cancelled) return;

    setMatches(body.data);
    setUpdatedAt(Date.now());
    setStatus("live");
    failures.current = 0;
  } catch (err) {
    if (cancelled || err.name === "AbortError") return;
    failures.current += 1;
    setStatus((prev) =&gt; (prev === "loading" ? "error" : "stale"));
  }
  schedule();
}

load();
return () =&gt; {
  cancelled = true;
  controller.abort();
  clearTimeout(timer);
};
</code></pre>
<p>  }, [sport, intervalMs]);</p>
<p>  return { matches, status, updatedAt };
}</p>
<p>What this small hook gets right:</p>
<p>Cleanup. When the component unmounts or the sport changes, the pending request is aborted and the timer cleared. Without this, old requests update new screens. This is the classic useEffect cleanup lesson.
Visibility check. If the tab is in the background, it skips polling. The browser exposes this through the Page Visibility API. It saves battery, data and your quota.
Backoff. If the network is failing, the hook slows down instead of hammering your server.
Four honest states. The user is never left guessing whether data is fresh, old or missing.
Built on fetch. No extra library needed.
Step 4: Designing the UI states</p>
<p>The difference between a polished app and a fragile one is mostly what you show when things go wrong.</p>
<p>jsx
export function LiveBoard({ sport }) {
  const { matches, status, updatedAt } = useLiveMatches(sport);</p>
<p>  if (status === "loading") return </p><p>Loading live matches...</p>;
  if (status === "error") return <p>Couldn't load matches. Retrying...</p>;
  if (matches.length === 0) return <p>No live matches right now. Check back soon.</p>;<p></p>
<p>  return (
    </p><section>
      {status === "stale" &amp;&amp; (
        <p>
          Connection problems. Showing data from{" "}
          {new Date(updatedAt).toLocaleTimeString()}.
        </p>
      )}
      <ul>
        {matches.map((m) =&gt; (
          
        ))}
      </ul>
    </section>
  );
}<p></p>
<p>Four states to design on purpose:</p>
<p>Loading. First visit, nothing yet.
Empty. The request worked, but there are no live matches. This is not an error and should not look like one.
Stale. You have old data and the network is failing. Show the old data plus a warning. Blanking the screen is worse.
Error. You have nothing to show.</p>
<p>Showing a "last updated" time in the stale state is a small touch that builds a lot of trust. Users forgive delays. They do not forgive being silently shown wrong scores.</p>
<p>Step 5: Making score changes noticeable</p>
<p>When a goal goes in, the score should feel like it changed. A short highlight does the job.</p>
<p>jsx
import { useEffect, useRef, useState } from "react";</p>
<p>function MatchRow({ match }) {
  const score = <code>${match.home.score} - ${match.away.score}</code>;
  const previous = useRef(score);
  const [flash, setFlash] = useState(false);</p>
<p>  useEffect(() =&gt; {
    if (previous.current !== score) {
      previous.current = score;
      setFlash(true);
      const t = setTimeout(() =&gt; setFlash(false), 2000);
      return () =&gt; clearTimeout(t);
    }
  }, [score]);</p>
<p>  return (
    &lt;li className={flash ? "row score-flash" : "row"}&gt;
      <span>{match.home.name}</span>
      <strong>{score}</strong>
      <span>{match.away.name}</span>
      <small>{match.minute ? <code>${match.minute}'</code> : match.status}</small>
    
  );
}
css
.score-flash { animation: pulse 2s ease-out; }
@keyframes pulse {
  0%   { background: #ffe9a8; }
  100% { background: transparent; }
}</p>
<p>Keep the animation short and subtle. Also make sure the change is not communicated by colour alone, since some users cannot see it. Text and position should carry the meaning too.</p>
<p>Step 6: Showing time correctly</p>
<p>Your proxy hands over kickoff as an ISO 8601 UTC string, like 2026-10-04T14:30:00Z. The browser knows the user's time zone, so let it do the work with Intl.DateTimeFormat:</p>
<p>javascript
export function formatKickoff(iso) {
  return new Intl.DateTimeFormat(undefined, {
    weekday: "short",
    hour: "2-digit",
    minute: "2-digit",
  }).format(new Date(iso));
}</p>
<p>formatKickoff("2026-10-04T14:30:00Z");   // "Sun, 8:00 PM" for a user in India</p>
<p>Passing undefined as the locale uses the user's own settings. You never have to guess their time zone, and you never store local times in your database.</p>
<p>Step 7: Upgrading from polling to streaming</p>
<p>Polling every 5 seconds is perfect for a first version. When you outgrow it, you can switch to a stream without rewriting the UI. The browser has a built-in EventSource for Server-Sent Events:</p>
<p>javascript
const source = new EventSource("/api/stream?sport=football");</p>
<p>source.onmessage = (event) =&gt; {
  const update = JSON.parse(event.data);
  // merge <code>update</code> into your React state
};</p>
<p>source.onerror = () =&gt; {
  // the browser reconnects on its own; show a "reconnecting" hint if you like
};</p>
<p>A clean approach is to keep polling as a fallback. If the stream fails to connect, the hook quietly drops back to polling, and the user never notices. Your UI components stay exactly the same, because they only care about matches, status and updatedAt.</p>
<p>Step 8: Offline and push notifications</p>
<p>Two features users love, both built on browser capabilities:</p>
<p>Offline support. A Service Worker can cache your app shell and the last known scores, so the app opens even with no signal. Show the "last updated" label so nobody mistakes cached data for live data.</p>
<p>Push notifications. The Push API lets your server notify users when something important happens, like a goal in a team they follow. The key rule is the same as always: deduplicate on the server so one goal produces one notification.</p>
<p>Both are optional for a first release. Get the board right first.</p>
<p>Multi-sport screens without the chaos</p>
<p>Football has minutes, cricket has overs, tennis has sets. If you build a separate screen for each sport, you will drown in code. A calmer approach:</p>
<p>Have the proxy return a common card shape (like toCard above).
Put sport-specific details in an optional detail field.
Render a generic row for the list view, and sport-specific layouts only on the match detail page.</p>
<p>This gets much easier when your data source already speaks one consistent format across sports, which is a big part of the idea behind Orbistats.</p>
<p>Common mistakes and quick fixes
Mistake	Fix
API key in frontend code	Move the call to a proxy, keep the key in an environment variable
Polling every second	Match the interval to how often the data really changes
Blank screen on error	Keep old data, show a "stale" notice
Request still running after navigation	Abort it in the effect cleanup
Polling in background tabs	Skip rounds when document.hidden
Showing kickoff in server time	Send UTC, format in the browser
Returning the entire provider payload	Trim it to what the UI uses
Trusting query parameters	Whitelist allowed values
No timeout on upstream calls	Always set one
A realistic build plan
Day 1: Proxy with one endpoint and a 3-second cache.
Day 2: React hook plus a basic list. Test it by killing your Wi-Fi.
Day 3: Loading, empty, stale and error states.
Day 4: Score highlight and correct time formatting.
Day 5: Add a second sport and see what needs generalising.
Later: Streaming, offline mode, notifications.</p>
<p>Testing the failure paths is the part most beginners skip, and it is the part that makes your app feel professional.</p>
<p>Final thoughts</p>
<p>A sports data API is only the beginning of the story. What reaches the user is shaped by dozens of small decisions in the middle: where the key lives, how long you cache, what you trim, how you behave when the network drops, how you show time, and how honest you are when data is old.</p>
<p>Get those right and even a simple 5-second polling app feels solid. Get them wrong and the best data feed in the world will still feel broken.</p>
<p>Start small, test the ugly cases, and grow from there. If you want to try a multi-sport data source while you build, take a look at Orbistats and tell me in the comments what you are making.</p>
<p>Disclosure: parts of this article were drafted with AI assistance and then reviewed and edited by the author.</p>
]]></content:encoded></item><item><title><![CDATA[A Beginner's Guide to Live Sports Data]]></title><description><![CDATA[Before the code, the data
Most tutorials on sports data jump straight into "here's how to call an API". That is useful, but it skips a more basic question: what does live sports data actually look lik]]></description><link>https://bettechmagnetics.hashnode.dev/a-beginner-s-guide-to-live-sports-data</link><guid isPermaLink="true">https://bettechmagnetics.hashnode.dev/a-beginner-s-guide-to-live-sports-data</guid><category><![CDATA[sports data]]></category><category><![CDATA[Python]]></category><category><![CDATA[data-engineering]]></category><category><![CDATA[Beginner Developers]]></category><category><![CDATA[SQLite]]></category><dc:creator><![CDATA[Vijay Choudhary]]></dc:creator><pubDate>Sat, 03 Oct 2026 19:38:46 GMT</pubDate><enclosure url="https://cdn.hashnode.com/uploads/covers/6a47ac8dec0b8a17b4dc915d/8ed5b50c-43d8-456a-8916-4b7e595c0a96.png" length="0" type="image/jpeg"/><content:encoded><![CDATA[<p>Before the code, the data</p>
<p>Most tutorials on sports data jump straight into "here's how to call an API". That is useful, but it skips a more basic question: what does live sports data actually look like, and how should you think about it?</p>
<p>If you do not understand the shape of the data, you will write code that works on a demo and breaks on a real match. A goal gets cancelled. A stat is corrected an hour later. Two sources disagree about the minute. Your database ends up with a score that does not match the TV.</p>
<p>This guide is for beginners who want to understand live sports data properly. We will cover:</p>
<p>The five kinds of live sports data
The difference between state and events (the most important idea here)
How different sports produce very different data
A clean event format you can reuse
How to store events in a small database
How to measure delay and analyse a match with Python
Free datasets to practise on before you pay for anything</p>
<p>You only need basic Python. No sports knowledge required.</p>
<p>What does "live" sports data mean?</p>
<p>"Live" does not mean instant. It means data that is produced while the match is still being played and delivered quickly enough to be useful.</p>
<p>That gap between the real-world moment and the moment the data reaches you is called latency, and every live source has some. If you want the formal idea, read about latency in engineering. For a fan app, a few seconds is usually fine. For anything time-sensitive, you will need to measure it yourself, and we will do exactly that later in this guide.</p>
<p>Live data also comes with a trade-off every beginner should know:</p>
<p>Fast, accurate, complete: you rarely get all three at the same moment.</p>
<p>The very first report of an event is fast but may be wrong. The verified version arrives later. Good systems are built to accept the fast version and then fix it.</p>
<p>The five kinds of live sports data</p>
<p>Almost everything you will meet falls into one of these groups.</p>
<ol>
<li><p>Match state. The current snapshot: score, clock, period, status (scheduled, live, finished, postponed). Small and simple.</p>
</li>
<li><p>Events. A list of things that happened: goal, foul, wicket, substitution, timeout. Each has a time, a team, often a player. Often called play-by-play or ball-by-ball data.</p>
</li>
<li><p>Statistics. Aggregated numbers: shots, possession, strike rate, rebounds. Many of these are derived from events. The classic summary table is the box score.</p>
</li>
<li><p>Tracking data. Positions of players and the ball, many times per second. Very detailed, very large, and usually sold at a premium.</p>
</li>
<li><p>Odds and probabilities. Prices from bookmakers, or model-based chances like expected goals. Powerful, but not a good first project.</p>
</li>
</ol>
<p>For your first build, stay with groups 1 to 3.</p>
<p>State vs events: the idea that changes everything</p>
<p>This is the most useful concept in the whole article.</p>
<p>State answers: "What is true right now?" (Team A leads 2-1 at 67')
Events answer: "What happened, in order?" (goal at 12', goal at 34', goal at 58')</p>
<p>Beginners usually store only state. Then something like this happens: the feed says "Team A 2-1", then a few minutes later "Team A 1-1". Why did it drop? You have no idea, because you never saved the history.</p>
<p>If you store events, you can always rebuild the state. You can also explain why it changed, fix mistakes, and replay the match later. State is easy to calculate. History is impossible to recover if you threw it away.</p>
<p>Rule: store the events, derive the state.</p>
<p>Every sport has a different shape</p>
<p>Same idea, different vocabulary. This is why a design built for one sport often cracks when you add a second.</p>
<p>Sport type	Time structure	Typical scoring event	Special cases
Football (soccer)	2 halves, running clock	Goal	Extra time, penalties, VAR reversals
Cricket	Innings, overs, balls	Runs, wickets	Rain rules like Duckworth-Lewis-Stern
Basketball	4 quarters, game clock	Points (1, 2, 3)	Overtime, timeouts
Tennis	Points, games, sets	Point	No clock at all
Baseball	Innings, outs	Run	Long-running statistical tradition, see sabermetrics
Motorsport	Laps	Position change	Pit stops, safety cars</p>
<p>Notice that tennis has no clock and motorsport has no "score". If you hardcode home_score, away_score and minute into your tables, those sports will not fit.</p>
<p>That is the challenge platforms like Orbistats are built around. It currently covers 13 sports ([add your 13 sports here]) and aims to give developers one consistent way to read data across them, so the sport-by-sport glue code shrinks.</p>
<p>A clean event format you can reuse</p>
<p>Here is a simple, sport-neutral event. The trick is to keep a few universal fields and push everything sport-specific into detail.</p>
<p>json
{
  "event_id": "evt_000451",
  "match_id": "m_90123",
  "seq": 451,
  "kind": "goal",
  "team": "home",
  "minute": 34,
  "occurred_at": "2026-10-04T14:12:08Z",
  "ref_event_id": null,
  "detail": { "player": "Player 9", "type": "header" }
}</p>
<p>A few small design choices, and why:</p>
<p>event_id is unique and stable, so a repeated message does not count twice.
seq is an increasing number, so you can spot gaps and out-of-order delivery.
ref_event_id lets a later event point at an earlier one. This is how a cancelled goal is represented: a new event that references the old one.
occurred_at is a UTC timestamp in ISO 8601 format. The Z means UTC. Always store UTC, and convert to local time only at the last moment, using Python's zoneinfo.
detail is a free-form object. Cricket can put over and ball in it. Tennis can put set and game.</p>
<p>Because the format is JSON, Python's built-in json module is all you need to read it.</p>
<p>Storing events in SQLite</p>
<p>You do not need a big database to start. SQLite ships with Python and runs from a single file.</p>
<p>python
import json
import sqlite3
from datetime import datetime, timezone</p>
<p>conn = sqlite3.connect("live.db")
conn.execute("""
CREATE TABLE IF NOT EXISTS events (
    event_id    TEXT PRIMARY KEY,
    match_id    TEXT NOT NULL,
    seq         INTEGER NOT NULL,
    kind        TEXT NOT NULL,
    team        TEXT,
    minute      INTEGER,
    occurred_at TEXT NOT NULL,
    received_at TEXT NOT NULL,
    raw         TEXT
)
""")
conn.execute("CREATE INDEX IF NOT EXISTS idx_events_match ON events(match_id, seq)")</p>
<p>def save_event(ev: dict) -&gt; None:
    received_at = datetime.now(timezone.utc).isoformat()
    conn.execute(
        """INSERT OR REPLACE INTO events
           (event_id, match_id, seq, kind, team, minute, occurred_at, received_at, raw)
           VALUES (?, ?, ?, ?, ?, ?, ?, ?, ?)""",
        (
            ev["event_id"], ev["match_id"], ev["seq"], ev["kind"],
            ev.get("team"), ev.get("minute"), ev["occurred_at"],
            received_at, json.dumps(ev),
        ),
    )
    conn.commit()</p>
<p>Three things worth noticing:</p>
<p>event_id is the primary key, and we use INSERT OR REPLACE. If the same event arrives twice, there is still only one row. If the provider sends a corrected version, it overwrites the old one.
We save received_at ourselves, separate from the provider's occurred_at. That gives us what we need to measure delay later.
We keep the raw message in raw. When something looks wrong next week, you will be glad you did.
Deriving the score from events</p>
<p>Now the "state from events" idea in code:</p>
<p>python
def current_score(match_id: str) -&gt; dict:
    rows = conn.execute(
        "SELECT event_id, kind, team, raw FROM events WHERE match_id = ? ORDER BY seq",
        (match_id,),
    ).fetchall()</p>
<pre><code>goals = {}
for event_id, kind, team, raw in rows:
    data = json.loads(raw)
    if kind == "goal":
        goals[event_id] = team
    elif kind == "goal_disallowed" and data.get("ref_event_id"):
        goals.pop(data["ref_event_id"], None)     # cancel the earlier goal

score = {"home": 0, "away": 0}
for team in goals.values():
    if team in score:
        score[team] += 1
return score
</code></pre>
<p>If a goal is later disallowed, you just add a goal_disallowed event pointing at it. Nothing is deleted, nothing is guessed, and the score corrects itself. That pattern matches how video review works in practice, for example football's VAR system.</p>
<p>Practise safely: build a match replay simulator</p>
<p>Waiting for a real match to test your code is slow and stressful. A better habit: replay an old match at high speed. Save a match as a .jsonl file (one JSON event per line) and feed it into your code.</p>
<p>python
import json
import time</p>
<p>def replay(path: str, speed: float = 60.0):
    """Replay events from a JSON Lines file. speed=60 means 1 match minute = 1 second."""
    with open(path, "r", encoding="utf-8") as f:
        events = [json.loads(line) for line in f if line.strip()]</p>
<pre><code>events.sort(key=lambda e: e["seq"])
previous_minute = 0

for ev in events:
    minute = ev.get("minute") or previous_minute
    wait = max(0, (minute - previous_minute) * 60 / speed)
    time.sleep(wait)
    previous_minute = minute
    yield ev
</code></pre>
<p>for ev in replay("sample_match.jsonl", speed=120):
    save_event(ev)
    print(ev["minute"], ev["kind"], current_score(ev["match_id"]))</p>
<p>Now you can test duplicates, out-of-order events and corrections in seconds. To simulate a duplicate, just call save_event(ev) twice and check that the score does not change.</p>
<p>Measuring delay like a grown-up</p>
<p>Earlier we said "live" is never instant. Let's measure it. We stored both when the provider says the event happened (occurred_at) and when we received it (received_at).</p>
<p>python
import sqlite3
import pandas as pd</p>
<p>conn = sqlite3.connect("live.db")
df = pd.read_sql_query(
    "SELECT * FROM events",
    conn,
    parse_dates=["occurred_at", "received_at"],
)</p>
<p>df["delay_s"] = (df["received_at"] - df["occurred_at"]).dt.total_seconds()</p>
<p>print(df["delay_s"].describe())
print(df["delay_s"].quantile([0.50, 0.95, 0.99]))</p>
<p>The pandas library does the heavy lifting. Look at the percentiles, not just the average. An average of 3 seconds can hide the fact that 1 in 100 events takes 20 seconds. The percentile idea is simple: p95 means "95% of events arrived faster than this number".</p>
<p>Two honest warnings:</p>
<p>This measures provider timestamp to your receive time, not the true real-world moment. It tells you how fast the feed reaches you, not how fast the provider noticed the event.
It is only meaningful if your computer's clock is accurate. Keep it synced.
A quick analysis: when are goals scored?</p>
<p>Once events are stored, analysis becomes easy.</p>
<p>python
goals = df[df["kind"] == "goal"].copy()
goals["window"] = (goals["minute"] // 15) * 15   # 0, 15, 30, 45, 60, 75</p>
<p>print(goals.groupby("window").size())</p>
<p>And a simple chart with Matplotlib:</p>
<p>python
import matplotlib.pyplot as plt</p>
<p>counts = goals.groupby("window").size()
counts.plot(kind="bar")
plt.xlabel("Match minute (start of 15-minute window)")
plt.ylabel("Goals")
plt.title("Goals by time window")
plt.tight_layout()
plt.show()</p>
<p>With one match this is just a toy. With a few hundred matches it starts telling real stories, which is exactly why storing clean events pays off.</p>
<p>Free data to practise with</p>
<p>You do not need to buy a feed to learn. These are well-known open resources:</p>
<p>StatsBomb open data for detailed football event data
Cricsheet for ball-by-ball cricket data
Retrosheet for baseball play-by-play history</p>
<p>Always read each project's licence and attribution terms before using the data in anything public or commercial.</p>
<p>When you are ready for genuinely live feeds, look at how professional providers document their products, for example SportsDataIO's developer portal, and compare it with what a multi-sport platform such as Orbistats offers.</p>
<p>Data quality: the part nobody puts in the tutorial</p>
<p>A short list of problems you should expect:</p>
<p>Corrections. Stats get revised after the match. Build for updates.</p>
<p>Duplicates. Networks retry. The same message can arrive twice.</p>
<p>Gaps. A missing seq number means you missed something. Ask the provider to resend, or re-fetch the match.</p>
<p>Out-of-order events. Event 12 can arrive before event 11. Sort by seq, not arrival time.</p>
<p>Postponed and abandoned matches. The status changes and the score may never finish.</p>
<p>Inconsistent names. "Man Utd", "Manchester United" and "Man United" are one team. Use IDs, not names.</p>
<p>Time zones. Store UTC. Always.</p>
<p>A word on responsible use</p>
<p>Sports data is licensed. Do not scrape a website and assume it is fine, because most sites forbid it in their terms. If you use a provider, read what you are allowed to do with the data: display it, store it, resell it, use it in a betting product. If your project touches gambling, check the rules in the places where your users live. A few minutes reading the terms now can save you a lot of trouble later.</p>
<p>A beginner roadmap</p>
<p>If you want a path from here, this one works:</p>
<p>Week 1: Download an open dataset and load it with pandas.
Week 2: Convert it into the event format above and store it in SQLite.
Week 3: Build the replay simulator and derive the score live.
Week 4: Add a second sport and see what breaks in your schema.
Then: Connect a real live source and measure its delay.</p>
<p>Step 4 is where most of the learning happens.</p>
<p>Final thoughts</p>
<p>Live sports data looks complicated from far away, but the core ideas are small: store events, derive state, always use UTC, expect duplicates and corrections, and measure instead of guessing.</p>
<p>Get those right with one sport and a tiny SQLite file, and you will find that adding the next sport, the next provider or the next thousand users is mostly an engineering problem you already know how to solve.</p>
<p>If you are curious what a multi-sport data platform looks like from the developer side, have a look at Orbistats. And tell me in the comments what you are planning to build.</p>
<p>Disclosure: parts of this article were drafted with AI assistance and then reviewed and edited by the author.</p>
]]></content:encoded></item><item><title><![CDATA[How Live Score Apps Get Their Data]]></title><description><![CDATA[The 3 seconds you never think about
You are watching a match. Your friend texts you "GOAL!!!" and you open your favourite score app. The score is already updated. A moment later, your phone buzzes wit]]></description><link>https://bettechmagnetics.hashnode.dev/how-live-score-apps-get-their-data</link><guid isPermaLink="true">https://bettechmagnetics.hashnode.dev/how-live-score-apps-get-their-data</guid><category><![CDATA[sports tech]]></category><category><![CDATA[api]]></category><category><![CDATA[Python]]></category><category><![CDATA[backend]]></category><category><![CDATA[Real Time]]></category><dc:creator><![CDATA[Vijay Choudhary]]></dc:creator><pubDate>Sat, 03 Oct 2026 19:32:21 GMT</pubDate><enclosure url="https://cdn.hashnode.com/uploads/covers/6a47ac8dec0b8a17b4dc915d/c224d14f-bd52-46bd-8f12-7d862ab1e210.png" length="0" type="image/jpeg"/><content:encoded><![CDATA[<p>The 3 seconds you never think about</p>
<p>You are watching a match. Your friend texts you "GOAL!!!" and you open your favourite score app. The score is already updated. A moment later, your phone buzzes with a push notification.</p>
<p>Somebody did not type that in. No intern is sitting at a desk refreshing a TV feed. Yet the number on your screen is right, and it arrived within seconds of the ball crossing the line.</p>
<p>So where did it come from? That is what this article is about. If you are a beginner developer curious about how this works, or you are planning to build a score app of your own, this is the full picture, step by step.</p>
<p>We will cover:</p>
<p>Where the raw data really originates
Who collects it and who sells it
How a provider's feed becomes your app's screen
How live updates travel to millions of phones
The corrections, delays and failures nobody talks about
Working code for the key pieces
Step 1: Where the data really starts</p>
<p>Here is the part that surprises most people: for many sports, the first source is a human being.</p>
<p>Data collection usually works in a few layers:</p>
<ol>
<li><p>Venue scouts and live data collectors. Trained operators watch the match, either in the stadium or via broadcast, and log events using specialised software: goal, foul, wicket, substitution, timeout. They are fast, because they have a keyboard layout built just for that sport.</p>
</li>
<li><p>Official league feeds. Many leagues run their own scoring systems. The scorer in the press box or the official scoreboard feeds an official source. Providers often license this directly.</p>
</li>
<li><p>Automated tracking technology. Camera and sensor systems measure things humans cannot. Ball-tracking systems like Hawk-Eye are used in several sports to track position and trajectory, which enables line calls and detailed stats.</p>
</li>
<li><p>Video review. Decisions can change after the fact. Football's Video Assistant Referee system is a perfect example: a goal can be scored, announced, and then reversed. Any serious data pipeline has to cope with that.</p>
</li>
</ol>
<p>A quick reality check: the fastest, most accurate score is usually the one with the most layers of verification behind it, and those layers cost time. That tension between fast and right shows up in every design decision later.</p>
<p>Step 2: The data providers in the middle</p>
<p>A small team building an app does not send scouts to stadiums. They buy the data from a provider.</p>
<p>Providers sit between the raw collection and your app. They clean, structure, standardise and deliver the data through APIs. Big names like Sportradar and SportsDataIO publish public docs, which are worth reading even if you never buy from them. Have a look at Sportradar's Sports Data API documentation and the SportsDataIO developer getting-started page to see what a professional feed looks like. For another style, Stats Perform's API catalogue shows how events, players, teams and standings are separated into different APIs.</p>
<p>What you are paying for is not only the data. It is also:</p>
<p>Speed: how quickly an event reaches you
Accuracy: how often it needs correcting
Consistency: the same structure for every match
Coverage: lower leagues and less popular sports included
Uptime: it must work on match day</p>
<p>Platforms like Orbistats sit in this layer too. Orbistats currently covers 13 sports ([add your 13 sports here]), delivered in one consistent structure so an app does not need a different integration per sport.</p>
<p>Step 3: How a feed reaches the app's backend</p>
<p>A provider offers one or more delivery methods. Your backend picks the one that fits.</p>
<p>REST polling. Your server asks every few seconds. Simple, good for schedules and standings, weaker for fast live events.</p>
<p>Webhooks. The provider calls your server when something happens. A webhook is just an HTTP request in the opposite direction.</p>
<p>Streaming (WebSockets or similar). One persistent connection, and events arrive the instant they exist. The protocol behind this is RFC 6455, and the browser-side API is documented on MDN.</p>
<p>Here is a simplified ingestion client for a streaming provider:</p>
<p>python
import asyncio
import json
import os
import websockets   # pip install websockets</p>
<p>API_KEY = os.environ["SPORTS_API_KEY"]
FEED_URL = "wss://stream.example.com/v1/live"</p>
<p>async def ingest(handle_raw_event):
    delay = 1
    while True:
        try:
            async with websockets.connect(
                FEED_URL,
                additional_headers={"Authorization": f"Bearer {API_KEY}"},
            ) as ws:
                delay = 1
                async for message in ws:
                    await handle_raw_event(json.loads(message))
        except (websockets.ConnectionClosed, OSError):
            await asyncio.sleep(delay)
            delay = min(delay * 2, 30)</p>
<p>The library used here is websockets. On older versions, the header argument is called extra_headers.</p>
<p>Step 4: Cleaning and normalising the data</p>
<p>Raw provider data is almost never ready to show to users. Team names differ, IDs differ, event types have different spellings, and every sport has a different structure.</p>
<p>So the backend runs a normaliser: a layer that converts everything into your app's own format. This is the single most valuable piece of code in a score app, because it protects the rest of your system from the outside world.</p>
<p>python
from dataclasses import dataclass, field
from datetime import datetime, timezone</p>
<p>@dataclass
class MatchEvent:
    event_id: str
    match_id: str
    seq: int
    kind: str                 # "goal", "wicket", "point", "goal_disallowed", ...
    team: str | None
    minute: int | None
    occurred_at: datetime
    ref_event_id: str | None = None   # used by corrections
    detail: dict = field(default_factory=dict)</p>
<p>def normalise_provider_a(raw: dict) -&gt; MatchEvent:
    return MatchEvent(
        event_id=str(raw["id"]),
        match_id=f'a:{raw["fixture"]}',
        seq=int(raw["sequence"]),
        kind=raw["type"].lower(),
        team=raw.get("team"),
        minute=raw.get("minute"),
        occurred_at=datetime.fromtimestamp(raw["ts"], tz=timezone.utc),
        ref_event_id=raw.get("ref"),
        detail=raw.get("extra", {}),
    )</p>
<p>Every provider you add gets its own small normalise_provider_x function. Everything downstream only ever sees MatchEvent. If you want your feed definitions to be self-documenting, specifying them with OpenAPI and JSON schemas is a good habit.</p>
<p>Step 5: Match state, and the art of handling corrections</p>
<p>A beginner mistake is to keep a single score number and add 1 whenever a goal arrives. That works until the first VAR decision.</p>
<p>The reliable pattern is to store the events and calculate the score from them:</p>
<p>python
def build_state(events: list[MatchEvent]) -&gt; dict:
    valid = {}
    for ev in sorted(events, key=lambda e: e.seq):
        if ev.kind == "goal":
            valid[ev.event_id] = ev
        elif ev.kind == "goal_disallowed" and ev.ref_event_id:
            valid.pop(ev.ref_event_id, None)        # reverse the goal</p>
<pre><code>score = {"home": 0, "away": 0}
for ev in valid.values():
    if ev.team in score:
        score[ev.team] += 1
return {"score": score, "goals": list(valid.values())}
</code></pre>
<p>Now a disallowed goal is just another event. The score fixes itself. Duplicates are harmless too: if the same event_id arrives twice, it overwrites instead of double counting.</p>
<p>Two more habits that save you on match day:</p>
<p>Ignore anything with a seq you have already processed (idempotency)
Keep the raw payload for debugging. When a user says "the score was wrong at 67 minutes", you will want proof of what you received
Step 6: Caching and the "fan-out" problem</p>
<p>Now the scary part. A big match can have hundreds of thousands of people watching at once. If every phone asked your database "what is the score?" every few seconds, you would fall over.</p>
<p>The answer is to separate reading from producing:</p>
<p>The backend computes match state once.
It stores the result in a fast in-memory store.
It broadcasts changes to everyone who is interested.</p>
<p>A common choice for the in-memory piece is Redis, whose publish/subscribe feature fits this perfectly. For bigger event pipelines, teams use a log such as Apache Kafka.</p>
<p>For data that does not change second by second (fixtures, standings, team lists), use plain HTTP caching. A CDN in front of your API serves the same response to thousands of users without touching your server. If you are new to this, Cloudflare's explanation of what a CDN does and MDN's HTTP caching guide are both excellent starting points.</p>
<p>A rule of thumb:</p>
<p>Data type	How to serve it
Teams, leagues, venues	Long cache (hours or days)
Fixtures, standings	Short cache (seconds to minutes)
Live score and events	Push, not poll
Step 7: Pushing live updates to phones and browsers</p>
<p>For live scores, the simplest option for most apps is Server-Sent Events (SSE). It is a one-way stream from server to client over normal HTTP, and browsers handle reconnection for you. See MDN's Server-sent events guide for the details.</p>
<p>Here is a minimal backend using FastAPI:</p>
<p>python
import asyncio
import json
from fastapi import FastAPI
from fastapi.responses import StreamingResponse</p>
<p>app = FastAPI()
subscribers: set[asyncio.Queue] = set()</p>
<p>async def publish(update: dict):
    for q in list(subscribers):
        if not q.full():          # slow client? skip instead of blocking everyone
            q.put_nowait(update)</p>
<p>@app.get("/stream")
async def stream():
    q: asyncio.Queue = asyncio.Queue(maxsize=100)
    subscribers.add(q)</p>
<pre><code>async def event_generator():
    try:
        while True:
            try:
                update = await asyncio.wait_for(q.get(), timeout=15)
                yield f"data: {json.dumps(update)}\n\n"
            except asyncio.TimeoutError:
                yield ": keep-alive\n\n"      # stops proxies closing the connection
    finally:
        subscribers.discard(q)

return StreamingResponse(event_generator(), media_type="text/event-stream")
</code></pre>
<p>And the browser side is almost embarrassingly short:</p>
<p>javascript
const source = new EventSource("/stream");</p>
<p>source.onmessage = (event) =&gt; {
  const update = JSON.parse(event.data);
  document.querySelector("#score").textContent =
    <code>${update.score.home} - ${update.score.away}</code>;
};</p>
<p>source.onerror = () =&gt; console.log("Connection lost, the browser will retry...");</p>
<p>When do you need WebSockets instead? When the client also needs to send messages back, for example subscribing to different matches on the fly. For a pure scoreboard, SSE is simpler and often enough.</p>
<p>Don't forget time. Every event should carry a UTC timestamp, and your servers should keep their clocks synced using NTP. If you ever want to measure how far behind the real world you are, you need trustworthy clocks.</p>
<p>Step 8: Push notifications</p>
<p>The buzz on your phone is a separate path. When the backend sees a "goal" event for a team you follow, it looks up who follows that team and hands a message to the mobile push services (Apple's and Google's). The important design point is deduplication: if the provider resends the same goal, your users must not get five notifications. This is exactly why the event_id and seq checks from earlier matter.</p>
<p>What actually goes wrong in production</p>
<p>Real-world score apps live with a long list of problems:</p>
<p>Delays. "Live" always has some lag. Different sources have different speeds, which is why two apps can show different scores for a few seconds.
Wrong first reports. An early scorer name or minute is sometimes revised later.
Reversals. VAR, review decisions, score changes after a protest.
Postponements and abandonments. The fixture simply changes status.
Provider outages. The feed drops mid-match.
Mismatched IDs. The same match has different IDs at different providers, so combining sources needs a mapping table.</p>
<p>The professional response to the outage problem is failover: keep a second source ready and switch when the first goes quiet. That only works if your normaliser gives both sources the same shape, which loops back to Step 4.</p>
<p>The multi-sport challenge</p>
<p>A football score app is one project. A 13-sport score app is a different job. Football counts minutes, cricket counts overs and wickets, tennis counts points within games within sets, basketball counts quarters. If your database was built around goals and minutes, adding a new sport can mean rewriting half the system.</p>
<p>Two ways to deal with it:</p>
<p>Design a generic event model (like the MatchEvent above) where sport-specific details live in a flexible detail field.
Use a provider that already gives one consistent structure across sports. This is the idea behind Orbistats, which gives you a consistent way to work with data across its 13 sports, so you write less glue code.
A simple architecture to remember
Stadium scouts / league feeds / tracking systems
                    |
                    v
          Data provider (cleans, structures)
                    |
                    v
   Your ingestion service (WebSocket / webhook / polling)
                    |
                    v
   Normaliser  -&gt;  Event store  -&gt;  Match state builder
                    |
                    v
        Redis cache / pub-sub  +  CDN for static data
                    |
                    v
       SSE / WebSocket fan-out  +  Push notifications
                    |
                    v
                 Your users</p>
<p>If you can draw this diagram from memory, you understand how most live score apps work.</p>
<p>Quick checklist before you build your own
Pick a provider and measure its delay yourself with real matches
Normalise everything into your own event format
Store events, derive the score from them
Make every handler idempotent
Use UTC everywhere
Cache static data, push live data
Plan for corrections and outages from day one
Test with a replayed old match before going live on a real one
Wrapping up</p>
<p>A live score looks like a tiny piece of text, but it is the end of a long relay: a scout or sensor, a provider, a feed, a normaliser, a cache, a stream, and finally your screen. Each step is simple on its own. The craft is in connecting them so that the whole thing stays fast, correct and calm when a match gets chaotic.</p>
<p>Build the small version first. One sport, one provider, one SSE stream. Then replay a match and see what breaks. That is where you learn the most.</p>
<p>If you are curious about a multi-sport data platform, check out Orbistats, and tell me in the comments what you are building.</p>
<p>Disclosure: parts of this article were drafted with AI assistance and then reviewed and edited by the author.</p>
]]></content:encoded></item><item><title><![CDATA[Sports Data APIs 101: Everything Beginners Need to Know]]></title><description><![CDATA[Why this guide exists
Almost every sports app you use has the same hidden engine underneath: a sports data API. The live score on your phone, the fantasy points that update during a match, the "win pr]]></description><link>https://bettechmagnetics.hashnode.dev/sports-data-apis-101-everything-beginners-need-to-know</link><guid isPermaLink="true">https://bettechmagnetics.hashnode.dev/sports-data-apis-101-everything-beginners-need-to-know</guid><category><![CDATA[Sports Technology]]></category><category><![CDATA[data-engineering]]></category><category><![CDATA[software architecture]]></category><category><![CDATA[api]]></category><category><![CDATA[distributed systems]]></category><dc:creator><![CDATA[Vijay Choudhary]]></dc:creator><pubDate>Sat, 03 Oct 2026 19:27:23 GMT</pubDate><enclosure url="https://cdn.hashnode.com/uploads/covers/6a47ac8dec0b8a17b4dc915d/df1440f2-521e-4272-b765-1df6d3c3652c.png" length="0" type="image/jpeg"/><content:encoded><![CDATA[<p>Why this guide exists</p>
<p>Almost every sports app you use has the same hidden engine underneath: a sports data API. The live score on your phone, the fantasy points that update during a match, the "win probability" bar on a streaming overlay. None of it is typed in by hand. Somewhere, a feed is pushing structured data, and an app is turning it into something you can read.</p>
<p>If you are a beginner, the first week with sports data is confusing. Everyone throws words around like "play-by-play", "delta updates", "rate limits" and "webhooks", and nobody stops to explain them. This guide fixes that.</p>
<p>By the end you will know:</p>
<p>What a sports data API actually is
What kind of data you can expect
How to make your first request in Python and JavaScript
How live data works (polling vs WebSockets vs SSE)
How to handle rate limits, retries and bad data
How to pick a provider without regretting it later</p>
<p>You do not need prior sports knowledge. If you can read basic Python or JavaScript, you are good.</p>
<p>What is a sports data API?</p>
<p>An API (Application Programming Interface) is a way for one program to ask another program for data. A sports data API is a service where you send a request like "give me today's live matches" and get back structured data, almost always in JSON.</p>
<p>Most sports APIs follow the standard request/response model of HTTP. You send a request to a URL (an endpoint), the server replies with a status code and a body. If you have ever opened a website, you have already used HTTP. An API just returns data instead of a web page.</p>
<p>Here is the simplest mental model:</p>
<p>Your app  ---&gt;  GET /matches?status=live  ---&gt;  Sports API
Your app  &lt;---  200 OK + JSON data         &lt;---  Sports API</p>
<p>That is 80% of it. The remaining 20% is authentication, speed, accuracy and scale, and that is where most beginner projects go wrong.</p>
<p>What data can you actually get?</p>
<p>Different providers offer different things, but most sports data falls into these buckets:</p>
<ol>
<li><p>Reference data (changes rarely)
Sports, leagues, seasons, teams, players, venues. Think of it as the dictionary of your app.</p>
</li>
<li><p>Schedules and fixtures
Which match happens when, where and between whom. Be careful with time zones here. We will come back to this.</p>
</li>
<li><p>Live scores and match state
Current score, match clock, period or inning, and status (scheduled, live, finished, postponed).</p>
</li>
<li><p>Play-by-play / event data
Every goal, wicket, foul, substitution or point as a separate event with a timestamp. This is the richest data and the hardest to handle well.</p>
</li>
<li><p>Statistics
Player and team stats, match stats, season totals, head-to-head records.</p>
</li>
<li><p>Standings and tables
League positions, points, net run rate, goal difference.</p>
</li>
<li><p>Odds
Prices from bookmakers for different markets. This needs extra care, because odds change quickly and different bookmakers name their markets differently. If you want to understand why odds are not the same as probabilities, the mathematics of bookmaking is a good starting point.</p>
</li>
</ol>
<p>A good beginner project only needs the first four. Do not start with odds.</p>
<p>REST, JSON and endpoints in 5 minutes</p>
<p>Most sports APIs are REST style. A few terms you will see everywhere:</p>
<p>Endpoint: a URL that returns a specific kind of data, e.g. /v1/matches
Query parameters: filters added to the URL, e.g. ?sport=cricket&amp;status=live
Headers: extra info sent with the request, like your API key
Status codes: the server's quick answer. 200 means OK, 401 means your key is wrong, 404 means not found, 429 means you are sending too many requests. The full list lives in the MDN status code reference.</p>
<p>A typical response looks like this:</p>
<p>json
{
  "data": [
    {
      "id": "m_90123",
      "sport": "football",
      "league": "Premier League",
      "status": "live",
      "minute": 67,
      "home": { "name": "Team A", "score": 2 },
      "away": { "name": "Team B", "score": 1 },
      "updated_at": "2026-10-04T14:32:10Z"
    }
  ]
}</p>
<p>Notice updated_at is in UTC (the Z at the end). Good APIs always do this. If an API gives you local times without a time zone, treat it as a red flag.</p>
<p>Authentication: your API key</p>
<p>Almost every provider gives you a key when you sign up. You send it with each request, usually in a header.</p>
<p>Golden rule: never hardcode your key in code you push to GitHub. Use environment variables.</p>
<p>bash
export SPORTS_API_KEY="your_key_here"
Your first request (Python)</p>
<p>Install the popular Requests library:</p>
<p>bash
pip install requests</p>
<p>Now fetch live matches:</p>
<p>python
import os
import requests</p>
<p>API_KEY = os.environ["SPORTS_API_KEY"]
BASE_URL = "<a href="https://api.example.com/v1">https://api.example.com/v1</a>"   # replace with your provider's URL</p>
<p>session = requests.Session()
session.headers.update({
    "Authorization": f"Bearer {API_KEY}",
    "Accept": "application/json",
})</p>
<p>def get_live_matches(sport: str):
    response = session.get(
        f"{BASE_URL}/matches",
        params={"sport": sport, "status": "live"},
        timeout=10,           # ALWAYS set a timeout
    )
    response.raise_for_status()
    return response.json()["data"]</p>
<p>for match in get_live_matches("football"):
    home, away = match["home"], match["away"]
    print(f'{home["name"]} {home["score"]} - {away["score"]} {away["name"]} ({match["minute"]}')')</p>
<p>Two beginner habits worth building right now: use a Session (it reuses connections and is faster), and always pass a timeout. Without one, your program can hang forever if the server stalls.</p>
<p>Same thing in JavaScript
javascript
const API_KEY = process.env.SPORTS_API_KEY;
const BASE_URL = "<a href="https://api.example.com/v1">https://api.example.com/v1</a>";</p>
<p>async function getLiveMatches(sport) {
  const url = new URL(<code>${BASE_URL}/matches</code>);
  url.searchParams.set("sport", sport);
  url.searchParams.set("status", "live");</p>
<p>  const res = await fetch(url, {
    headers: {
      Authorization: <code>Bearer ${API_KEY}</code>,
      Accept: "application/json",
    },
    signal: AbortSignal.timeout(10_000),
  });</p>
<p>  if (!res.ok) throw new Error(<code>API error: ${res.status}</code>);
  const body = await res.json();
  return body.data;
}</p>
<p>getLiveMatches("basketball").then(console.log).catch(console.error);</p>
<p>Before writing code against any API, try it in a tool like Postman first. Seeing the raw response saves hours of confusion.</p>
<p>Live data: three ways to get updates</p>
<p>This is the part that separates a toy project from a real one. A match changes every few seconds. How does your app find out?</p>
<ol>
<li>Polling (simplest)</li>
</ol>
<p>Your app asks again and again: "anything new?" every N seconds.</p>
<p>Pros: easy, works everywhere.
Cons: wasteful, and you only see changes at your polling interval.</p>
<p>A smarter way to poll is to use conditional requests. The server sends an ETag with the response. Next time you send it back, and if nothing changed the server replies 304 Not Modified with no body, which saves bandwidth (and sometimes your quota).</p>
<p>python
import time</p>
<p>url = f"{BASE_URL}/matches/m_90123"
etag = None</p>
<p>while True:
    headers = {"If-None-Match": etag} if etag else {}
    r = session.get(url, headers=headers, timeout=10)</p>
<pre><code>if r.status_code == 304:
    pass                         # nothing changed, do nothing
elif r.ok:
    etag = r.headers.get("ETag")
    print("Updated:", r.json()["data"]["home"]["score"])
else:
    print("Error:", r.status_code)

time.sleep(5)
</code></pre>
<p>Not every provider supports ETags, so check the docs.</p>
<ol>
<li>Server-Sent Events (one-way stream)</li>
</ol>
<p>With Server-Sent Events, the server keeps one HTTP connection open and pushes updates to you. It is one-directional (server to client), simple, and reconnects automatically in browsers. A nice middle ground if your provider offers it.</p>
<ol>
<li>WebSockets (two-way stream)</li>
</ol>
<p>WebSockets, defined in RFC 6455, give you a persistent two-way connection. You subscribe once, and the provider pushes events the moment they happen. This is what serious live products use.</p>
<p>python
import asyncio
import json
import os
import websockets   # pip install websockets</p>
<p>API_KEY = os.environ["SPORTS_API_KEY"]
WS_URL = "wss://stream.example.com/v1/live"</p>
<p>async def listen(sport: str):
    last_seq = 0           # last event number we processed
    delay = 1              # reconnect delay (seconds)</p>
<pre><code>while True:
    try:
        # note: older versions of the library call this extra_headers
        async with websockets.connect(
            WS_URL,
            additional_headers={"Authorization": f"Bearer {API_KEY}"},
        ) as ws:
            await ws.send(json.dumps({
                "action": "subscribe",
                "sport": sport,
                "since": last_seq,      # ask to resume where we left off
            }))
            delay = 1                   # connected fine, reset the delay

            async for raw in ws:
                event = json.loads(raw)

                if event["seq"] &lt;= last_seq:
                    continue            # duplicate or old event, skip it

                last_seq = event["seq"]
                print(event["type"], event["match_id"], event.get("detail"))

    except (websockets.ConnectionClosed, OSError):
        print(f"Disconnected. Retrying in {delay}s...")
        await asyncio.sleep(delay)
        delay = min(delay * 2, 30)      # exponential backoff, capped
</code></pre>
<p>asyncio.run(listen("cricket"))</p>
<p>Three things in this snippet are what real pipelines do:</p>
<p>Sequence numbers: so you can detect duplicates and gaps
Resume from the last seen event: so a dropped connection does not lose data
Exponential backoff: so you do not hammer the server while it is struggling</p>
<p>Which one should you pick? Quick guide:</p>
<p>Need	Pick
Prototype, personal project	Polling
Live updates in a browser, one-way	Server-Sent Events
Real-time product, many matches at once	WebSockets
Rate limits: the thing that will surprise you</p>
<p>Every API limits how many requests you can send. Go over and you get 429 Too Many Requests. Good providers also send a Retry-After header telling you how long to wait. The semantics are defined in RFC 9110.</p>
<p>Here is a small retry helper you can reuse:</p>
<p>python
import time
import random
import requests</p>
<p>def get_with_retry(session, url, max_tries=5, **kwargs):
    for attempt in range(max_tries):
        r = session.get(url, timeout=10, **kwargs)</p>
<pre><code>    if r.status_code == 429:
        wait = int(r.headers.get("Retry-After", 2 ** attempt))
    elif r.status_code &gt;= 500:
        wait = 2 ** attempt + random.random()   # a little jitter
    else:
        r.raise_for_status()
        return r

    time.sleep(wait)

raise RuntimeError(f"Gave up after {max_tries} tries: {url}")
</code></pre>
<p>Practical ways to stay under limits:</p>
<p>Cache reference data (teams, leagues, venues). It barely changes.
Don't poll finished matches. Stop when status is finished.
Fetch lists, not single items, when you need many matches.
Prefer streaming once you pass a handful of live matches.
The data problems nobody warns you about</p>
<p>Pulling data is easy. Trusting it is the hard part. These are the issues that bite beginners.</p>
<p>Time zones. Store everything in UTC. Convert to local time only when displaying. Kick-off times in the wrong time zone are the number one beginner bug in sports apps.</p>
<p>IDs differ between providers. Provider A calls a match m_90123, provider B calls the same match 8841002. If you ever combine two sources, you need a mapping table.</p>
<p>Corrections happen. A goal gets disallowed after a VAR review. A scorer is changed. A stat is corrected the next day. Your app must be able to update or remove an event, not just add new ones.</p>
<p>Delays. "Live" rarely means instant. Every provider has some latency between the real event and the data reaching you. For a fan app a few seconds is fine. For anything time-sensitive, measure it.</p>
<p>Different sports, different shapes. A football match has goals and minutes. A cricket match has overs, wickets and innings. Tennis has sets, games and points. If your database assumes one shape, it will break the moment you add a second sport.</p>
<p>Working with many sports</p>
<p>This last point deserves its own section, because it is where projects quietly become painful.</p>
<p>Supporting one sport is a feature. Supporting ten is an architecture problem. Each sport has its own scoring rules, its own period structure and its own vocabulary. Many teams end up stitching together a different provider per sport, each with a different format, authentication style and ID scheme.</p>
<p>A better approach is to look for a provider that gives you one consistent structure across sports. That is the idea behind Orbistats, which currently covers 13 sports: [add your 13 sports here, e.g. Football, Cricket, Basketball, ...]. When the response shape stays predictable from one sport to the next, you write your parsing code once instead of thirteen times.</p>
<p>Whatever provider you evaluate, ask for a unified schema, because it will save you more time than any single feature.</p>
<p>How to choose a sports data provider</p>
<p>Before you commit, run through this checklist:</p>
<p>Coverage: Does it cover the sports and leagues you need, including the smaller ones?
Freshness: How fast does live data arrive, and can they show you measured numbers?
Delivery options: REST only, or also WebSockets/streaming?
Documentation quality: Is there an OpenAPI spec you can import into your tools?
Rate limits and pricing: What happens when you grow 10x?
Historical data: Do you need past seasons for analytics or ML?
Free tier or trial: Can you test with real data before paying?
Support: Is a human reachable when a feed breaks on match day?</p>
<p>It also helps to look at how the big names present themselves, so you know what "standard" looks like. The SportsDataIO developer portal and Sportradar's Sports Data API docs are both good examples of structured documentation. Reading two or three providers' docs side by side teaches you the common patterns fast.</p>
<p>Common beginner mistakes (and quick fixes)
Mistake	Fix
Polling every second "to be safe"	Respect the docs, use ETags or streaming
No timeout on requests	Always set one
Storing local times	Store UTC
Trusting the first score update	Handle corrections and reversals
Hardcoding the API key	Use environment variables
One database schema per sport	Design a flexible event model
Ignoring 429 responses	Back off and retry using Retry-After
Five beginner projects to try
Live scoreboard in the terminal using the polling script above
Discord or Telegram bot that posts goals as they happen
Personal dashboard showing your favourite team's next 5 fixtures
Standings tracker that stores the table daily and charts the movement
Multi-sport ticker that displays live matches from different sports in one clean feed</p>
<p>The last one is the best learning project, because it forces you to design a data model that works across sports.</p>
<p>Final thoughts</p>
<p>Sports data APIs look simple from the outside: send a request, get JSON. The real skill is in everything around that: handling live updates properly, surviving rate limits, trusting (and correcting) the data, and keeping your design flexible enough to add another sport without a rewrite.</p>
<p>Start small. Build the polling scoreboard. Then move to WebSockets. Then add a second sport and see what breaks. Every problem you hit is one that production systems face too, just at bigger scale.</p>
<p>If you want to see what a multi-sport data platform looks like in practice, take a look at Orbistats, and let me know in the comments what you are building.</p>
<p>Disclosure: parts of this article were drafted with AI assistance and then reviewed and edited by the author.</p>
]]></content:encoded></item><item><title><![CDATA[Microservices vs Monolith: Architecting a Sports Betting Backend
]]></title><description><![CDATA[Every team building a betting product eventually has the same argument. One side wants microservices from day one, because odds, wallets and settlement "obviously" scale differently. The other side wa]]></description><link>https://bettechmagnetics.hashnode.dev/microservices-vs-monolith-architecting-a-sports-betting-backend</link><guid isPermaLink="true">https://bettechmagnetics.hashnode.dev/microservices-vs-monolith-architecting-a-sports-betting-backend</guid><category><![CDATA[System Design]]></category><category><![CDATA[Microservices]]></category><category><![CDATA[PostgreSQL]]></category><category><![CDATA[architecture]]></category><category><![CDATA[Node.js]]></category><dc:creator><![CDATA[Vijay Choudhary]]></dc:creator><pubDate>Tue, 29 Sep 2026 13:46:15 GMT</pubDate><enclosure url="https://cdn.hashnode.com/uploads/covers/6a47ac8dec0b8a17b4dc915d/c3a1e069-71c2-4a8e-8851-c852642eb4c7.png" length="0" type="image/jpeg"/><content:encoded><![CDATA[<p>Every team building a betting product eventually has the same argument. One side wants microservices from day one, because odds, wallets and settlement "obviously" scale differently. The other side wants a monolith, because they have six engineers and a launch date. Both sides are right about something, and both are usually wrong about when.</p>
<p>This guide takes a position and then defends it with code: start with a modular monolith, keep money and bets in one transactional core, and extract the parts whose scaling and failure profile genuinely differ, starting with odds ingestion and the streaming layer. Along the way we will build the parts that are easy to get wrong in a betting system: a double-entry ledger, idempotent bet placement, a fail-closed price feed, settlement, and an outbox for reliable events.</p>
<p>Read this first</p>
<p>This is an architecture walkthrough, not a benchmark and not legal advice. I use Orbistats as the example upstream data provider because its public docs describe REST, WebSocket and Webhooks delivery for fixtures, live scores, statistics and odds, and its homepage now advertises 13 sports. I make no claims about how Orbistats or any operator builds its internals. Field names in the code are placeholders, so compare real payloads in the public API sandbox and change only the adapter layer.</p>
<p>Three cautions apply to everything below:</p>
<p>A data feed is not your product. Odds arriving from a provider are an input to your trading and pricing decisions. They are not automatically the prices you are allowed or willing to offer customers.
Betting is regulated. Licensing, age and identity verification, geographic restrictions, anti-money-laundering controls and responsible-gambling tooling shape the architecture as much as scale does. Part 12 covers this.
Numbers in code are illustrative. Timeouts, staleness windows and limits are defaults to tune against your own traffic and your regulator's rules.
Part 1: What a betting backend actually does</p>
<p>Strip away the branding and a sportsbook backend is a set of capabilities that change at very different speeds:</p>
<p>Capability	What it does	Rate of change	Consistency needs
Data ingestion	Fixtures, live state, events, odds from providers	Continuous, bursty	Eventual is fine
Pricing and trading	Turns feed data into offered prices	Continuous	Must be fresh, fail closed
Streaming to clients	Pushes prices and scores to apps	Continuous, high fan-out	Eventual, lossy is acceptable
Bet placement	Validates price, funds, limits, then commits a bet	Per customer action	Strong, transactional
Wallet and ledger	Records every movement of money	Per money event	Strong, auditable, append-only
Risk and exposure	Tracks liability per market	Per bet	Strong enough to enforce limits
Settlement	Resolves bets when results are official	Bursty at match end	Idempotent, correctable
Compliance	KYC state, geo, self-exclusion, limits	Rare	Strong at bet time
Notifications and webhooks	Tells users and partners things happened	Per event	Eventual, retried</p>
<p>Look at the last column. The rows that need strong consistency (placement, wallet, risk, compliance) all touch the same money. The rows that tolerate eventual consistency (ingestion, streaming, notifications) touch none of it. That split is the whole architecture in one sentence, and it is why the monolith-versus-microservices question has a more useful answer than "it depends".</p>
<p>Part 2: The monolith and the microservices, honestly compared</p>
<p>A monolith is one deployable unit. Modules call each other in-process, and a single database transaction can span all of them.</p>
<p>Microservices are independently deployable services, each owning its data, communicating over the network.</p>
<p>Concern	Monolith	Microservices
Money consistency	One ACID transaction across wallet, bet, risk	Sagas, compensations, eventual consistency
Development speed (small team)	Fast	Slower: contracts, deployments, tooling
Independent scaling	Scale everything together	Scale the hot service only
Fault isolation	One bad module can take everything down	Failures can be contained
Deployment risk	Every change ships the whole thing	Small, independent deploys
Operational cost	One thing to run and observe	Many things, plus the network between them
Team autonomy	Shared codebase, needs discipline	Natural ownership boundaries
Debugging	One stack trace	Distributed tracing required
Data ownership	Shared database, discipline needed	Enforced by the network</p>
<p>Two honest observations:</p>
<p>The hardest problem in betting, moving money correctly, gets much harder when you split it across a network. A bet that debits a wallet in one service and creates a bet record in another needs a saga, idempotent retries and reconciliation. All of that is solvable, and none of it is free.
The problem microservices solve best, independent scaling of a spiky, stateless workload, exists in betting but only in a few places, mainly ingestion and streaming.
Part 3: The recommendation, a modular monolith with a planned exit</p>
<p>A modular monolith has the deployment simplicity of a monolith and the internal boundaries of services. Each module owns its tables and exposes a small public interface. Other modules may only call that interface, never reach into internals or read foreign tables.</p>
<p>The payoff is optionality. When a module genuinely needs its own process, you extract it along a boundary that already exists. When it never does, you never paid for it.</p>
<p>src/
  modules/
    odds/            # ingestion, feed prices, offered prices, price stream
    catalog/         # sports, competitions, fixtures, markets, selections
    wallet/          # accounts, ledger, deposits and withdrawals (interfaces)
    bets/            # placement, cash-out, settlement
    risk/            # limits, exposure, stake rules
    compliance/      # KYC status, geo checks, self-exclusion, deposit limits
    notifications/   # push, email, outbound webhooks
  platform/
    db.ts            # connection pool, transaction helper
    bus.ts           # event bus interface and implementations
    outbox.ts        # reliable event publishing
    config.ts
  app.ts</p>
<p>The rule that makes this work is mechanical, so enforce it in CI instead of in code review. With dependency-cruiser, deep imports across modules fail the build:</p>
<p>js
// .dependency-cruiser.cjs
module.exports = {
  forbidden: [
    {
      name: 'only-import-module-public-api',
      comment: 'Modules may import another module only through its index.ts',
      severity: 'error',
      from: { path: '^src/modules/([^/]+)/' },
      to: {
        path: '^src/modules/[^/]+/(?!index\.ts$)',
        pathNot: '^src/modules/$1/',
      },
    },
  ],
};</p>
<p>Add a second convention: a module never queries another module's tables. If bets needs a price, it calls odds.getOfferedPrice(). If a boundary is respected inside the monolith, extracting it later means replacing a function call with a network call, not untangling a database.</p>
<p>Part 4: The internal contract, an event bus you can swap</p>
<p>Modules need to talk asynchronously too. Define the bus as an interface from day one:</p>
<p>ts
// platform/bus.ts
export interface DomainEvent&lt;T = unknown&gt; {
  id: string;            // unique, used by consumers to deduplicate
  type: string;          // 'odds.changed', 'bet.placed', 'fixture.finished'
  occurredAt: string;    // ISO 8601, UTC
  data: T;
}</p>
<p>export type Handler = (e: DomainEvent) =&gt; Promise;</p>
<p>export interface EventBus {
  publish(e: DomainEvent): Promise;
  subscribe(type: string, handler: Handler): void;
}</p>
<p>// In-process implementation for the monolith
export class InProcessBus implements EventBus {
  private handlers = new Map&lt;string, Handler[]&gt;();</p>
<p>  subscribe(type: string, h: Handler) {
    this.handlers.set(type, [...(this.handlers.get(type) ?? []), h]);
  }</p>
<p>  async publish(e: DomainEvent) {
    for (const h of this.handlers.get(e.type) ?? []) {
      try { await h(e); }
      catch (err) { console.error('handler failed', e.type, e.id, err); }
    }
  }
}</p>
<p>Later you can add a RedisStreamsBus, NatsBus or KafkaBus implementing the same interface, and modules will not notice. One warning: an in-process bus loses events if the process crashes between commit and publish. For events that matter to money or customer trust, use the outbox in Part 8.</p>
<p>Part 5: The ingestion side, from feed to offered price</p>
<p>The upstream provides the raw material. Orbistats describes a Sports Data API for fixtures and reference data, a Live Scores API for match state and events, a Sports Statistics API, an Odds API with normalized pre-match and live markets, and a Historical Sports Data API useful for backtesting your own risk and pricing models. For streaming, see the WebSocket API and Webhooks API. Check the documentation and changelog for exact payloads and recent changes.</p>
<p>Keep two price tables with very different meanings:</p>
<p>sql
-- What the provider says. Input only. Never shown to customers directly.
CREATE TABLE feed_prices (
  selection_id bigint  NOT NULL,
  source       text    NOT NULL,           -- 'orbistats'
  price        numeric(9,3) NOT NULL CHECK (price &gt; 1),
  suspended    boolean NOT NULL DEFAULT false,
  source_ts    timestamptz NOT NULL,
  received_at  timestamptz NOT NULL DEFAULT now(),
  PRIMARY KEY (selection_id, source)
);</p>
<p>-- What YOU are willing to offer. Set by pricing/trading logic and risk.
CREATE TABLE offered_prices (
  selection_id bigint PRIMARY KEY,
  price        numeric(9,3) NOT NULL CHECK (price &gt; 1),
  suspended    boolean NOT NULL DEFAULT true,     -- fail closed by default
  updated_at   timestamptz NOT NULL DEFAULT now()
);</p>
<p>The gap between those two tables is where your trading decisions live: margin, limits by market, manual overrides, and suspension rules. Whether and how you may derive offered prices from a third-party feed is a licensing question (Part 12), so settle it before you build.</p>
<p>The ingestion module applies feed updates with an ordering guard, so a delayed old message cannot overwrite a newer one:</p>
<p>ts
// modules/odds/ingest.ts
import type { Pool } from 'pg';
import type { EventBus } from '../../platform/bus';</p>
<p>export async function applyFeedPrice(
  pool: Pool,
  bus: EventBus,
  u: { selectionId: number; price: number; suspended: boolean; sourceTs: Date },
) {
  const res = await pool.query(
    <code>INSERT INTO feed_prices (selection_id, source, price, suspended, source_ts)      VALUES ($1, 'orbistats', $2, $3, $4)      ON CONFLICT (selection_id, source) DO UPDATE        SET price = EXCLUDED.price,            suspended = EXCLUDED.suspended,            source_ts = EXCLUDED.source_ts,            received_at = now()      WHERE feed_prices.source_ts &lt;= EXCLUDED.source_ts      RETURNING selection_id</code>,
    [u.selectionId, u.price, u.suspended, u.sourceTs],
  );</p>
<p>  if (res.rowCount) {
    await bus.publish({
      id: crypto.randomUUID(),
      type: 'odds.feed_changed',
      occurredAt: new Date().toISOString(),
      data: { selectionId: u.selectionId },
    });
  }
}</p>
<p>A pricing subscriber listens for odds.feed_changed, applies your margin and risk rules, and upserts offered_prices. Keep that logic behind odds.getOfferedPrice(selectionId) so nothing else depends on the tables.</p>
<p>Fail closed when the feed goes quiet</p>
<p>In a betting system, stale prices are not a cosmetic problem. They are a way to lose money and to accept bets you should not. If the feed goes silent, suspend.</p>
<p>ts
// modules/odds/watchdog.ts
import type { Pool } from 'pg';</p>
<p>let lastFeedMessageAt = Date.now();
export const markFeedAlive = () =&gt; { lastFeedMessageAt = Date.now(); };</p>
<p>const STALE_AFTER_MS = 5000;   // illustrative: tune to your feed's real cadence</p>
<p>export function startWatchdog(pool: Pool) {
  setInterval(async () =&gt; {
    if (Date.now() - lastFeedMessageAt &lt;= STALE_AFTER_MS) return;
    await pool.query(
      <code>UPDATE offered_prices op        SET suspended = true, updated_at = now()        FROM selections s        JOIN markets m  ON m.id = s.market_id        JOIN fixtures f ON f.id = m.fixture_id        WHERE op.selection_id = s.id          AND f.status IN ('live', 'paused')          AND op.suspended = false</code>,
    );
  }, 1000);
}</p>
<p>Call markFeedAlive() in the upstream WebSocket message handler. Use reconnect with exponential backoff and jitter, as in the 50-line live scoreboard walkthrough, and copy the exact subscribe payload from the docs.</p>
<p>Part 6: The wallet, a double-entry ledger</p>
<p>Never store "the balance" as a number you mutate freely. Store movements, and keep the balance as a cache that the database itself keeps honest.</p>
<p>sql
CREATE TABLE accounts (
  id            bigint GENERATED ALWAYS AS IDENTITY PRIMARY KEY,
  owner_user_id uuid,                          -- NULL for system accounts
  kind          text NOT NULL CHECK (kind IN
                  ('user_cash','bets_held','house_pnl','payments_clearing')),
  currency      char(3) NOT NULL,
  balance_minor bigint NOT NULL DEFAULT 0,     -- minor units: cents, paise
  UNIQUE NULLS NOT DISTINCT (owner_user_id, kind, currency),
  CONSTRAINT no_negative_user_cash
    CHECK (kind &lt;&gt; 'user_cash' OR balance_minor &gt;= 0)
);</p>
<p>CREATE TABLE ledger_transactions (
  id         bigint GENERATED ALWAYS AS IDENTITY PRIMARY KEY,
  ref_type   text NOT NULL,                    -- 'bet_place','bet_settle','deposit'...
  ref_id     text NOT NULL,
  created_at timestamptz NOT NULL DEFAULT now(),
  UNIQUE (ref_type, ref_id)                    -- the same event cannot post twice
);</p>
<p>CREATE TABLE ledger_entries (
  id          bigint GENERATED ALWAYS AS IDENTITY PRIMARY KEY,
  tx_id       bigint NOT NULL REFERENCES ledger_transactions(id),
  account_id  bigint NOT NULL REFERENCES accounts(id),
  delta_minor bigint NOT NULL CHECK (delta_minor &lt;&gt; 0)
);
CREATE INDEX ledger_entries_account_idx ON ledger_entries (account_id, id);</p>
<p>Three properties do the heavy lifting:</p>
<p>Every transaction sums to zero. Money is never created or destroyed by a bug, only moved.
UNIQUE (ref_type, ref_id) makes posting idempotent. Replaying a settlement job cannot pay a bet twice.
A CHECK constraint on user_cash means even a buggy code path cannot overdraw a customer. The database refuses.</p>
<p>The posting helper enforces the zero-sum rule and updates accounts in a fixed order to avoid deadlocks:</p>
<p>ts
// modules/wallet/ledger.ts
import type { PoolClient } from 'pg';</p>
<p>export interface Entry { accountId: number; deltaMinor: number; }</p>
<p>export async function postLedger(
  db: PoolClient,
  ref: { type: string; id: string },
  entries: Entry[],
): Promise {
  const sum = entries.reduce((s, e) =&gt; s + e.deltaMinor, 0);
  if (sum !== 0) throw new Error(<code>unbalanced ledger transaction (${sum})</code>);</p>
<p>  const tx = await db.query(
    <code>INSERT INTO ledger_transactions (ref_type, ref_id) VALUES ($1, $2) RETURNING id</code>,
    [ref.type, ref.id],
  );
  const txId = Number(tx.rows[0].id);</p>
<p>  // Sorted by account id: every transaction locks rows in the same order
  const ordered = [...entries].sort((a, b) =&gt; a.accountId - b.accountId);
  for (const e of ordered) {
    await db.query(
      <code>INSERT INTO ledger_entries (tx_id, account_id, delta_minor) VALUES ($1, $2, $3)</code>,
      [txId, e.accountId, e.deltaMinor],
    );
    await db.query(
      <code>UPDATE accounts SET balance_minor = balance_minor + $2 WHERE id = $1</code>,
      [e.accountId, e.deltaMinor],
    );
  }
  return txId;
}</p>
<p>Notes for production: use integer minor units everywhere (never floats), treat pg bigint values as strings or BigInt once amounts can exceed 2^53, and be aware that shared system accounts like bets_held and house_pnl can become hot rows. Shard them per currency and per hour, or compute their balances from entries instead of updating a single row.</p>
<p>Part 7: Bet placement, the one transaction that must be right</p>
<p>This is the strongest argument for the monolith. Placing a bet needs to atomically:</p>
<p>Confirm the customer is allowed to bet (compliance)
Confirm the price is still valid and fresh (odds)
Confirm the fixture and market are open (catalog)
Confirm exposure and stake limits (risk)
Move the stake from the customer's cash to a held account (wallet)
Record the bet (bets)
Queue the bet.placed event (outbox)</p>
<p>In a monolith that is one database transaction. In a microservice design that is a saga with five failure points.</p>
<p>sql
CREATE TABLE bets (
  id                     uuid PRIMARY KEY DEFAULT gen_random_uuid(),
  user_id                uuid NOT NULL,
  idempotency_key        text NOT NULL,
  selection_id           bigint NOT NULL,
  fixture_id             bigint NOT NULL,
  stake_minor            bigint NOT NULL CHECK (stake_minor &gt; 0),
  currency               char(3) NOT NULL,
  accepted_price         numeric(9,3) NOT NULL CHECK (accepted_price &gt; 1),
  potential_payout_minor bigint NOT NULL,
  status                 text NOT NULL DEFAULT 'open'
                         CHECK (status IN ('open','won','lost','void','cashed_out')),
  placed_at              timestamptz NOT NULL DEFAULT now(),
  settled_at             timestamptz,
  UNIQUE (user_id, idempotency_key)
);
CREATE INDEX bets_open_by_selection ON bets (selection_id) WHERE status = 'open';</p>
<p>The client generates an idempotencyKey per bet attempt. If the network drops after the server commits, the retry returns the same bet instead of creating a second one.</p>
<p>ts
// modules/bets/place.ts
import type { Pool } from 'pg';
import { postLedger } from '../wallet';
import { compliance } from '../compliance';
import { risk } from '../risk';</p>
<p>export type PriceChangePolicy = 'reject' | 'accept_better' | 'accept_any';</p>
<p>export interface PlaceBetCommand {
  userId: string;
  idempotencyKey: string;
  selectionId: number;
  stakeMinor: number;
  currency: string;
  expectedPrice: string;           // the decimal price the customer saw
  priceChangePolicy: PriceChangePolicy;
}</p>
<p>export class BetError extends Error {
  constructor(public code: string, public detail?: unknown) { super(code); }
}</p>
<p>const MAX_PRICE_AGE_MS = 3000;     // illustrative</p>
<p>export async function placeBet(pool: Pool, cmd: PlaceBetCommand) {
  // Cheap, read-only gate first (KYC, age, geo, self-exclusion, deposit limits)
  await compliance.assertCanBet(cmd.userId);</p>
<p>  const db = await pool.connect();
  try {
    await db.query('BEGIN');</p>
<pre><code>// 1. Serialize all activity for this customer by locking their wallet
const wallet = await db.query(
  `SELECT id, balance_minor FROM accounts
   WHERE owner_user_id = $1 AND kind = 'user_cash' AND currency = $2
   FOR UPDATE`,
  [cmd.userId, cmd.currency],
);
if (!wallet.rowCount) throw new BetError('no_wallet');

// 2. Idempotency: a retry returns the original result
const prior = await db.query(
  `SELECT id, status, accepted_price FROM bets
   WHERE user_id = $1 AND idempotency_key = $2`,
  [cmd.userId, cmd.idempotencyKey],
);
if (prior.rowCount) { await db.query('COMMIT'); return prior.rows[0]; }

// 3. Current offered price and market state, read inside the transaction
const q = await db.query(
  `SELECT op.price, op.suspended, op.updated_at,
          f.id AS fixture_id, f.status AS fixture_status
   FROM offered_prices op
   JOIN selections s ON s.id = op.selection_id
   JOIN markets m    ON m.id = s.market_id
   JOIN fixtures f   ON f.id = m.fixture_id
   WHERE op.selection_id = $1`,
  [cmd.selectionId],
);
if (!q.rowCount) throw new BetError('unknown_selection');
const p = q.rows[0];

if (p.suspended)                                  throw new BetError('market_suspended');
if (!['scheduled', 'live'].includes(p.fixture_status)) throw new BetError('fixture_closed');
if (Date.now() - new Date(p.updated_at).getTime() &gt; MAX_PRICE_AGE_MS)
                                                  throw new BetError('price_stale');

// 4. Price-change policy: what the customer agreed to accept
const offered = Number(p.price);
const expected = Number(cmd.expectedPrice);
const changed = offered !== expected;
const worse = offered &lt; expected;
if (changed &amp;&amp; (cmd.priceChangePolicy === 'reject' ||
    (worse &amp;&amp; cmd.priceChangePolicy === 'accept_better'))) {
  throw new BetError('price_changed', { offered: p.price });
}

// 5. Funds (the CHECK constraint is the backstop, this gives a clean error)
if (Number(wallet.rows[0].balance_minor) &lt; cmd.stakeMinor) throw new BetError('insufficient_funds');

// 6. Risk: stake limits and exposure, reserved inside THIS transaction
await risk.reserveExposure(db, {
  selectionId: cmd.selectionId,
  stakeMinor: cmd.stakeMinor,
  price: p.price,
});

// 7. Record the bet; payout is computed in SQL to avoid float rounding
const ins = await db.query(
  `INSERT INTO bets (user_id, idempotency_key, selection_id, fixture_id, stake_minor,
                     currency, accepted_price, potential_payout_minor)
   VALUES ($1, $2, $3, $4, $5, $6, $7, floor($5::numeric * $7::numeric)::bigint)
   RETURNING id, status, accepted_price, potential_payout_minor`,
  [cmd.userId, cmd.idempotencyKey, cmd.selectionId, p.fixture_id,
   cmd.stakeMinor, cmd.currency, p.price],
);
const bet = ins.rows[0];

// 8. Move the stake: user cash -&gt; bets held
const held = await db.query(
  `SELECT id FROM accounts WHERE kind = 'bets_held' AND currency = $1`, [cmd.currency]);
await postLedger(db, { type: 'bet_place', id: bet.id }, [
  { accountId: Number(wallet.rows[0].id), deltaMinor: -cmd.stakeMinor },
  { accountId: Number(held.rows[0].id),   deltaMinor:  cmd.stakeMinor },
]);

// 9. Reliable event, committed atomically with everything above
await db.query(
  `INSERT INTO outbox (type, payload) VALUES ('bet.placed', $1)`,
  [JSON.stringify({ betId: bet.id, userId: cmd.userId, selectionId: cmd.selectionId })],
);

await db.query('COMMIT');
return bet;
</code></pre>
<p>  } catch (err) {
    await db.query('ROLLBACK');
    throw err;
  } finally {
    db.release();
  }
}</p>
<p>Look at what the monolith bought us. The exposure reservation, the wallet debit, the bet row and the event all commit or roll back together. There is no window in which the customer is charged without a bet, or has a bet without a charge. In a split design you rebuild that guarantee with sagas, timeouts, reconciliation jobs and manual repair tooling.</p>
<p>Also note what the design leaves as choices: many operators add a deliberate acceptance delay for live betting so the price can be re-checked, and some place bets through an asynchronous pending state. Those are product and risk decisions. The important part is that whichever you choose, the money movement stays in one transaction.</p>
<p>Part 8: The outbox, reliable events without a distributed transaction</p>
<p>Publishing to a bus after COMMIT can fail, and publishing before can announce something that then rolls back. The outbox pattern solves both: write the event in the same transaction as the state change, and let a relay publish it.</p>
<p>sql
CREATE TABLE outbox (
  id           bigint GENERATED ALWAYS AS IDENTITY PRIMARY KEY,
  type         text NOT NULL,
  payload      jsonb NOT NULL,
  created_at   timestamptz NOT NULL DEFAULT now(),
  published_at timestamptz
);
CREATE INDEX outbox_unpublished ON outbox (id) WHERE published_at IS NULL;
ts
// platform/outbox.ts
import type { Pool } from 'pg';
import type { EventBus } from './bus';</p>
<p>export async function relayOutbox(pool: Pool, bus: EventBus, batch = 100): Promise {
  const db = await pool.connect();
  try {
    await db.query('BEGIN');
    const { rows } = await db.query(
      <code>SELECT id, type, payload FROM outbox        WHERE published_at IS NULL        ORDER BY id        LIMIT $1        FOR UPDATE SKIP LOCKED</code>,
      [batch],
    );
    for (const r of rows) {
      await bus.publish({
        id: String(r.id),
        type: r.type,
        occurredAt: new Date().toISOString(),
        data: r.payload,
      });
    }
    if (rows.length) {
      await db.query(<code>UPDATE outbox SET published_at = now() WHERE id = ANY($1)</code>,
        [rows.map(r =&gt; r.id)]);
    }
    await db.query('COMMIT');
    return rows.length;
  } catch (e) {
    await db.query('ROLLBACK');
    throw e;
  } finally {
    db.release();
  }
}</p>
<p>FOR UPDATE SKIP LOCKED lets several relay workers run without stepping on each other. Delivery is at least once, so every consumer deduplicates by event id. That is not a flaw to hide. It is the contract, and it is exactly the same contract you will have after you move to an external broker.</p>
<p>Part 9: Settlement, idempotent and correctable</p>
<p>Settlement resolves open bets when a result becomes official. It has two enemies: running twice, and being wrong.</p>
<p>ts
// modules/bets/settle.ts
import type { Pool } from 'pg';
import { postLedger } from '../wallet';</p>
<p>export type Outcome = 'won' | 'lost' | 'void';</p>
<p>export async function settleSelection(pool: Pool, selectionId: number, outcome: Outcome) {
  for (;;) {
    const db = await pool.connect();
    try {
      await db.query('BEGIN');</p>
<pre><code>  // Claim a batch of still-open bets; concurrent workers skip locked rows
  const { rows: bets } = await db.query(
    `SELECT id, user_id, currency, stake_minor, potential_payout_minor
     FROM bets
     WHERE selection_id = $1 AND status = 'open'
     ORDER BY id
     LIMIT 200
     FOR UPDATE SKIP LOCKED`,
    [selectionId],
  );
  if (!bets.length) { await db.query('COMMIT'); return; }

  for (const b of bets) {
    const stake = Number(b.stake_minor);
    const payout = Number(b.potential_payout_minor);

    const acc = await db.query(
      `SELECT
         (SELECT id FROM accounts WHERE owner_user_id = $1 AND kind = 'user_cash' AND currency = $2) AS cash,
         (SELECT id FROM accounts WHERE kind = 'bets_held' AND currency = $2) AS held,
         (SELECT id FROM accounts WHERE kind = 'house_pnl' AND currency = $2) AS house`,
      [b.user_id, b.currency],
    );
    const { cash, held, house } = acc.rows[0];

    const entries =
      outcome === 'won'
        ? [ { accountId: Number(held),  deltaMinor: -stake },
            { accountId: Number(cash),  deltaMinor:  payout },
            { accountId: Number(house), deltaMinor: -(payout - stake) } ]
      : outcome === 'lost'
        ? [ { accountId: Number(held),  deltaMinor: -stake },
            { accountId: Number(house), deltaMinor:  stake } ]
        : [ { accountId: Number(held),  deltaMinor: -stake },
            { accountId: Number(cash),  deltaMinor:  stake } ];

    await postLedger(db, { type: 'bet_settle', id: b.id }, entries);
    await db.query(
      `UPDATE bets SET status = $2, settled_at = now() WHERE id = $1 AND status = 'open'`,
      [b.id, outcome],
    );
    await db.query(
      `INSERT INTO outbox (type, payload) VALUES ('bet.settled', $1)`,
      [JSON.stringify({ betId: b.id, userId: b.user_id, outcome })],
    );
  }
  await db.query('COMMIT');
} catch (e) {
  await db.query('ROLLBACK');
  throw e;
} finally {
  db.release();
}
</code></pre>
<p>  }
}</p>
<p>Why this is safe to re-run: the ledger's UNIQUE (ref_type, ref_id) rejects a duplicate posting, and the status = 'open' filter means settled bets are never picked up again. A crashed job can simply start over.</p>
<p>Sport rules belong in a rules layer, not in this function</p>
<p>The function above settles a selection. Deciding which outcome a selection has is where sports differ, and it is where most settlement incidents come from. With 13 sports on the Orbistats homepage, each with its own page such as tennis, plan for rules like these, and confirm exact coverage per sport in the docs before you offer a market:</p>
<p>Football and Handball: which period counts (regular time or including extra time), abandoned matches
Tennis and Volleyball: retirements and walkovers before or after a set is complete
Cricket: rain-shortened matches and DLS outcomes, ties and super overs
Baseball: listed-pitcher rules and games shortened by weather
Golf: ties and dead-heat rules, players who miss the cut or withdraw
Horse Racing: non-runners, dead heats and place-terms changes
Combat Sports: draws, no contests and stoppages
Esports: map counts, forfeits and rescheduled series
Ice Hockey and Basketball: overtime inclusion
American Football: overtime and postponed games</p>
<p>Put these in a settlement-rules package keyed by sport slug, with one small pure function per market type. Pure functions are cheap to unit test against recorded results.</p>
<p>Results get corrected, so version them</p>
<p>Providers revise data: a goal is reassigned, a result changes after a protest. Do not hard-code "settled means final". Store a result_version with each settlement decision, keep an audit of what triggered it, and build a deliberate resettlement path that reverses the original ledger transaction with a compensating one instead of editing history. Read the provider's changelog and agree internally on who is authorised to trigger a resettlement.</p>
<p>Part 10: When to extract a service, and in what order</p>
<p>A service earns its existence when one of these is true. Otherwise leave it in the monolith.</p>
<p>Signal	Example	Why it justifies a split
Different scaling shape	Streaming to a large number of connected clients	Scales with connections, not with bets
Different failure blast radius	Feed ingestion crashing should not stop bet placement	Isolation is worth the network hop
Different resource profile	CPU-heavy pricing or simulation jobs	Keeps noisy work off the transactional core
Different release cadence	Notifications or widgets changing weekly	Independent deploys without touching money code
Different compliance scope	Storing sensitive personal data	Smaller audit surface, tighter access
A team boundary	A separate trading team owning pricing	Ownership matches the service</p>
<p>Applying that test to the modules we built gives an extraction order:</p>
<p>Odds ingestion and the price stream (first). It is stateless enough to restart freely, bursty, and its failure mode (go stale, suspend) is already designed. It talks to the core through the bus and offered_prices, so the seam already exists.
Notifications and outbound webhooks. Purely event-driven, tolerant of delay, and they benefit from independent retry policies.
Settlement workers. Same codebase and database, deployed as a separate worker process, so heavy settlement bursts at full time do not compete with bet placement. This is a deployment split, not a data split.
Pricing and trading logic, if a separate team emerges.
Wallet, bets, risk and compliance stay together as the core for as long as you can defend it. Splitting them is the expensive one. Do it only when you have concrete evidence that this transactional core is your scaling bottleneck, and be ready to invest in sagas, reconciliation and idempotency at every hop.
Extraction is a swap, not a rewrite</p>
<p>Because odds is behind an interface, the extraction is mechanical:</p>
<p>ts
// modules/odds/index.ts (public API of the module)
export interface PriceReader {
  getOfferedPrice(selectionId: number): Promise&lt;{ price: string; suspended: boolean; updatedAt: Date }&gt;;
}</p>
<p>Day one, that reads offered_prices in the same database. After extraction, the odds service owns its own database and publishes odds.offered_changed events, and the core keeps a local read-model updated by those events. Bet placement then reads the local read-model, so the hot path never depends on a network call to another service at the moment money moves. That is the pattern to hold onto: services may publish state, but the transactional core validates against data it already owns.</p>
<p>Part 11: Failure modes and the response you designed in
Failure	What could go wrong	Design response
Feed WebSocket drops	Stale prices accepted	Watchdog suspends live prices, reconnect with backoff and jitter
Malformed feed payload	Bad price offered	Validate at the adapter, reject and alert, never coerce
Duplicate bet request	Double charge	Idempotency key, unique constraint, wallet row lock
Two bets race on one wallet	Overdraft	FOR UPDATE on the wallet row plus CHECK (balance &gt;= 0)
Crash after commit, before publish	Lost event	Transactional outbox
Settlement job runs twice	Double payout	Unique ledger reference plus status guard
Result corrected after settlement	Wrong payouts	Versioned results, compensating ledger entries
Hot system account	Lock contention at full time	Shard system accounts or derive balances from entries
Slow streaming client	Memory growth	Buffer limit, drop, then disconnect
Provider changes an API	Sudden breakage	Adapter isolation, changelog monitoring, contract tests on saved sandbox responses</p>
<p>The last row deserves a habit: save real responses from the sandbox, replay them in CI, and fail the build when a mapper breaks. Replay a recorded match day with duplicated and reordered messages, and assert the end state equals the ordered run.</p>
<p>Part 12: Licensing, compliance and responsible gambling shape the architecture</p>
<p>This part is not optional, and it is not legal advice. Consult qualified counsel and your regulator for your jurisdictions.</p>
<p>Data licensing. Check what your provider agreement allows: storage duration, redistribution, using feed prices as inputs to your own pricing, and commercial use in a wagering product. Orbistats positions enterprise plans around custom feeds, dedicated infrastructure and SLA requirements, so raise these questions early. Review the pricing page for plan details, and read the About page and developer hub to understand how the provider positions itself. Its sub-50ms feed-latency figure is the vendor's own marketing statement, so measure your own end-to-end path and never promise customers a speed you have not measured.</p>
<p>Operator compliance. Depending on where you operate, expect requirements such as:</p>
<p>Identity and age verification before wagering, with re-verification triggers
Geographic restrictions enforced at bet time, not only at login
Anti-money-laundering monitoring on deposits, withdrawals and unusual patterns
Self-exclusion, cooling-off periods, and customer-set deposit, loss and time limits that must be enforced server-side
Immutable audit logs and long, regulated retention periods
Complaints handling and data protection obligations</p>
<p>Architecturally, that means the compliance module sits in the bet placement path, its checks run on every bet, and its decisions are logged. It also means a strong reason to keep it inside the transactional core initially: a limit that is enforced eventually is a limit that can be breached.</p>
<p>Part 13: Observability from the first day</p>
<p>Whichever shape you choose, instrument the seams:</p>
<p>Bet placement: success rate, rejection reasons by code (price_changed, price_stale, market_suspended, insufficient_funds), and p95 and p99 latency
Feed health: time since last message, reconnect count, suspension events
Ledger: an hourly job asserting that all transaction deltas sum to zero and that account balances equal the sum of their entries
Outbox: age of the oldest unpublished row
Settlement: bets settled per minute, resettlements, failures by sport
Per-tenant or per-partner labels if you expose an API, as in a multi-tenant design</p>
<p>The reconciliation job is the one people skip and later regret:</p>
<p>sql
-- Should return zero rows. Any row is an incident.
SELECT a.id, a.balance_minor, coalesce(sum(e.delta_minor), 0) AS ledger_sum
FROM accounts a
LEFT JOIN ledger_entries e ON e.account_id = a.id
GROUP BY a.id, a.balance_minor
HAVING a.balance_minor &lt;&gt; coalesce(sum(e.delta_minor), 0);
Part 14: A phased plan</p>
<p>Phase 1, one deployable, clean modules. Modules catalog, odds, wallet, bets, risk, compliance. Ledger, idempotent placement, outbox, in-process bus, CI-enforced boundaries.</p>
<p>Phase 2, separate the workers. Same code, run ingestion, the outbox relay and settlement as separate processes so their load does not compete with the API. Add the feed watchdog and reconciliation job.</p>
<p>Phase 3, extract ingestion and streaming. Put a real broker behind the EventBus interface, move the WebSocket fan-out to its own service, and keep a local read-model of offered prices in the core. Add the Webhooks API as an alternative inbound path where event-driven delivery fits better than a persistent socket.</p>
<p>Phase 4, extract by evidence. Split pricing, notifications or anything else only when metrics show a real scaling, isolation or ownership reason.</p>
<p>Phase 5, revisit the core. Only if the transactional core is a demonstrated bottleneck after sharding hot accounts, tuning indexes, adding read replicas and partitioning historical tables, do you consider splitting money-critical paths. Budget for sagas and reconciliation before you start.</p>
<p>If you are exploring what the data side looks like first, build the smallest client and study real payloads: the live scoreboard walkthrough, the Python FastAPI wrapper and the .NET Core client are quick ways to see the shapes you will later map in your adapter.</p>
<p>Common mistakes
Splitting into microservices before you have a second team or a measured bottleneck.
Treating the odds feed price as the offered price.
Mutating a balance column with no ledger behind it.
Placing a bet across two services without a real answer for partial failure.
Trusting an in-process event bus for events that involve money.
Failing open when the feed goes quiet.
Using floats for odds or money.
Hard-coding "final result" with no way to correct and resettle.
Enforcing responsible-gambling limits eventually instead of at bet time.
Calling a shared database with many services a microservice architecture. That is a distributed monolith with extra latency.
A repeatable checklist
Separate what needs strong consistency (money, bets, risk, compliance) from what tolerates eventual consistency (ingestion, streaming, notifications).
Start with a modular monolith and enforce module boundaries in CI.
Define the event bus as an interface before you need a broker.
Keep feed_prices and offered_prices separate, and fail closed on stale data.
Use a double-entry ledger with unique references and database-level balance checks.
Make bet placement one idempotent transaction with a price-change policy.
Publish events through a transactional outbox and deduplicate on the consumer.
Settle idempotently, keep sport rules in pure functions, and version results.
Extract ingestion and streaming first, and keep the money core together until evidence says otherwise.
Treat licensing, compliance and responsible gambling as architecture inputs.
Final thoughts</p>
<p>The monolith-versus-microservices debate is really a question about where consistency matters and where scale differs. In a betting backend, consistency matters most exactly where the money moves, and scale differs most exactly where data flows in and prices flow out. Follow that seam. Keep the transactional heart in one place, put a clean interface around everything else, and let real metrics tell you which module deserves to leave home first.</p>
<p>If you have extracted a service from a betting or trading system and regretted it (or been glad you did), tell me which one in the comments. Those are the stories that actually change how other teams decide.</p>
]]></content:encoded></item><item><title><![CDATA[The Complete Guide to Sports Data Schema Design for Startups]]></title><description><![CDATA[The Complete Guide to Sports Data Schema Design for Startups
Every sports startup begins the same way. You call an API, get a friendly JSON blob, and save it in a table with columns like home_team, aw]]></description><link>https://bettechmagnetics.hashnode.dev/the-complete-guide-to-sports-data-schema-design-for-startups</link><guid isPermaLink="true">https://bettechmagnetics.hashnode.dev/the-complete-guide-to-sports-data-schema-design-for-startups</guid><category><![CDATA[PostgreSQL]]></category><category><![CDATA[database]]></category><category><![CDATA[System Design]]></category><category><![CDATA[api]]></category><category><![CDATA[TypeScript]]></category><dc:creator><![CDATA[Vijay Choudhary]]></dc:creator><pubDate>Tue, 29 Sep 2026 13:42:34 GMT</pubDate><enclosure url="https://cdn.hashnode.com/uploads/covers/6a47ac8dec0b8a17b4dc915d/175486e1-043a-4014-895e-daa767250b73.png" length="0" type="image/jpeg"/><content:encoded><![CDATA[<p>The Complete Guide to Sports Data Schema Design for Startups</p>
<p>Every sports startup begins the same way. You call an API, get a friendly JSON blob, and save it in a table with columns like home_team, away_team, home_score and away_score. It works for one league. Then someone asks for tennis. Then cricket. Then golf, where there is no home team at all, and horse racing, where a single event has fourteen participants.</p>
<p>This guide is about designing the database so that moment is a Tuesday afternoon task, not a rewrite. We will build a schema in Postgres, map an upstream feed into it with TypeScript, and talk through the mistakes that are cheap to avoid on day one and painful to fix in year two.</p>
<p>About this guide</p>
<p>This is a design walkthrough, not a benchmark and not a description of any vendor's internal database. I use Orbistats as the example upstream because its public docs describe fixtures, results, standings, teams, players, events, statistics and odds across many sports, and because its homepage now advertises 13 sports. The schema itself works with any licensed feed. Field names in the mapping code are placeholders, so compare real payloads in the public API sandbox and change only the adapter.</p>
<p>Who this is for
A founder or first backend engineer building a scores, stats, fantasy, media or analytics product
A team that started with football and knows more sports are coming
Anyone who has ever written if (sport === 'tennis') inside a database query
Part 1: Principles before tables</p>
<p>Seven rules shape every decision below.</p>
<p>Your internal IDs are yours. Never use a provider's ID as your primary key. Providers change, merge and re-issue identifiers, and you may add a second provider.
Model the shape of competition, not the sport. Football, tennis and horse racing differ in rules, but they share a structure: an event, some competitors, scored periods, a sequence of things that happened, and a final result.
Append what happened, derive what is true now. Events and price ticks are facts. Current score and current odds are conclusions you can recompute.
Store the boring things as columns and the strange things as JSONB. Columns for what you filter and join on. JSONB for sport-specific details you only display.
Make ingestion idempotent. Feeds retry, reorder and repeat. Writing the same update twice must be harmless.
Keep time in UTC and lifecycle in an explicit state. Never infer "finished" from a timestamp.
Keep the raw payload for a while. When a mapping bug appears, replaying stored input is the difference between an hour and a week.
Part 2: What the upstream gives you</p>
<p>Before designing tables, list the raw materials. The Orbistats documentation describes a sport, then resource, then endpoint model, with resources such as fixtures, fixture details, results, standings, odds, statistics, lineups, events, teams, players, competitions and countries. The product pages split into the Sports Data API for schedules and reference data, the Live Scores API for match state and events, the Sports Statistics API, the Odds API, and the Historical Sports Data API for archives. Delivery is via REST, the WebSocket API and the Webhooks API.</p>
<p>That gives us the entity list: sports, countries, competitions, seasons, competitors (teams or individuals), players, venues, fixtures, periods, events, lineups, statistics, standings, and odds. Now we give each one a home.</p>
<p>Part 3: The entity map
sports ──&lt; competitions ──&lt; seasons ──&lt; fixtures ──&lt; fixture_competitors &gt;── competitors ──&lt; competitor_members &gt;── players
                                            │
                                            ├──&lt; fixture_periods
                                            ├──&lt; fixture_results
                                            ├──&lt; fixture_events
                                            ├──&lt; competitor_fixture_stats / player_fixture_stats
                                            └──&lt; markets ──&lt; selections ──&lt; odds_current / odds_ticks</p>
<p>seasons ──&lt; standings_rows          external_ids (provider mapping)         raw_ingest (replayable input)</p>
<p>The single most important design choice is in the middle: fixtures connect to competitors through a slot-based join table, not through home_id and away_id columns. That one decision is what lets the same schema hold a football match, a tennis doubles match, a golf tournament and a horse race.</p>
<p>Part 4: Reference tables</p>
<p>Start with the sports themselves. Seeding them as data, not as an enum in code, lets you add a sport with an INSERT.</p>
<p>sql
CREATE TABLE sports (
  id          smallint PRIMARY KEY,
  slug        text NOT NULL UNIQUE,      -- matches the API path segment
  code        text NOT NULL UNIQUE,      -- short code shown on the website
  name        text NOT NULL,
  structure   text NOT NULL CHECK (structure IN ('head_to_head', 'field', 'race')),
  period_kind text NOT NULL              -- default label: half, quarter, set, innings, round, map...
);</p>
<p>INSERT INTO sports (id, slug, code, name, structure, period_kind) VALUES
  (1,  'football',          'FTB', 'Football',          'head_to_head', 'half'),
  (2,  'basketball',        'BSK', 'Basketball',        'head_to_head', 'quarter'),
  (3,  'american-football', 'NFL', 'American Football', 'head_to_head', 'quarter'),
  (4,  'cricket',           'CRK', 'Cricket',           'head_to_head', 'innings'),
  (5,  'tennis',            'TEN', 'Tennis',            'head_to_head', 'set'),
  (6,  'baseball',          'BSB', 'Baseball',          'head_to_head', 'inning'),
  (7,  'esports',           'ESP', 'Esports',           'head_to_head', 'map'),
  (8,  'combat-sports',     'MMA', 'Combat Sports',     'head_to_head', 'round'),
  (9,  'volleyball',        'VBL', 'Volleyball',        'head_to_head', 'set'),
  (10, 'handball',          'HBL', 'Handball',          'head_to_head', 'half'),
  (11, 'ice-hockey',        'ICE', 'Ice Hockey',        'head_to_head', 'period'),
  (12, 'golf',              'GLF', 'Golf',              'field',        'round'),
  (13, 'horse-racing',      'HRC', 'Horse Racing',      'race',         'race')
ON CONFLICT (id) DO NOTHING;</p>
<p>Those 13 rows mirror the sport cards on the Orbistats homepage, each with its own page such as tennis. Confirm the exact path slugs for the sports that are not in the documentation examples against the official documentation, and re-check before every release because coverage can differ by product page.</p>
<p>Now countries, competitions and seasons.</p>
<p>sql
CREATE TABLE countries (
  code char(2) PRIMARY KEY,
  name text NOT NULL
);</p>
<p>CREATE TABLE competitions (
  id            bigint GENERATED ALWAYS AS IDENTITY PRIMARY KEY,
  sport_id      smallint NOT NULL REFERENCES sports(id),
  country_code  char(2) REFERENCES countries(code),   -- NULL for international
  name          text NOT NULL,
  tier          smallint,
  attributes    jsonb NOT NULL DEFAULT '{}'
);</p>
<p>CREATE TABLE seasons (
  id             bigint GENERATED ALWAYS AS IDENTITY PRIMARY KEY,
  competition_id bigint NOT NULL REFERENCES competitions(id),
  label          text NOT NULL,          -- '2026/27' or '2026'
  starts_on      date,
  ends_on        date,
  UNIQUE (competition_id, label)
);</p>
<p>seasons is not optional decoration. Standings, fixtures and statistics all belong to a season, and without it you cannot answer "how did this team do last year" without date arithmetic that breaks on split seasons.</p>
<p>Part 5: Competitors, players and membership</p>
<p>Here is the first place naive schemas fail. Football has teams. Tennis has players. Doubles has pairs. Golf has individuals. Cricket has teams whose members change per match. If you create teams and players and then bolt player_a_id onto fixtures, tennis works and everything else hurts.</p>
<p>The fix is one level of indirection: a competitor is whatever appears on a scoresheet. It might be a club, a person, a doubles pair, or a horse and jockey entry.</p>
<p>sql
CREATE TABLE players (
  id          bigint GENERATED ALWAYS AS IDENTITY PRIMARY KEY,
  name        text NOT NULL,
  birth_date  date,
  country_code char(2) REFERENCES countries(code),
  attributes  jsonb NOT NULL DEFAULT '{}'
);</p>
<p>CREATE TABLE competitors (
  id           bigint GENERATED ALWAYS AS IDENTITY PRIMARY KEY,
  sport_id     smallint NOT NULL REFERENCES sports(id),
  kind         text NOT NULL CHECK (kind IN ('team', 'individual', 'pair', 'entry')),
  name         text NOT NULL,
  short_name   text,
  country_code char(2) REFERENCES countries(code),
  attributes   jsonb NOT NULL DEFAULT '{}'
);</p>
<p>CREATE TABLE competitor_members (
  competitor_id bigint NOT NULL REFERENCES competitors(id),
  player_id     bigint NOT NULL REFERENCES players(id),
  role          text,                    -- 'captain', 'jockey', 'partner' ...
  valid_from    date NOT NULL DEFAULT '1900-01-01',
  valid_to      date,
  PRIMARY KEY (competitor_id, player_id, valid_from)
);</p>
<p>For a singles tennis player, you create one individual competitor with one member. That feels like overhead for the first week and pays for itself the first time you add doubles, relays or racing entries. If you are truly football-only for a year, you can collapse this later, but going the other direction is much harder. The upstream's stable team and player identifiers, which its solution pages describe for multi-season analysis, map directly onto external_ids, which we cover in Part 9.</p>
<p>Part 6: Fixtures without home and away</p>
<p>A fixture is a scheduled contest. Notice what is not in this table: scores, team IDs, and anything sport-specific.</p>
<p>sql
CREATE TABLE venues (
  id       bigint GENERATED ALWAYS AS IDENTITY PRIMARY KEY,
  name     text NOT NULL,
  city     text,
  timezone text,                          -- IANA, e.g. 'Asia/Kolkata'
  country_code char(2) REFERENCES countries(code)
);</p>
<p>CREATE TABLE fixtures (
  id                bigint GENERATED ALWAYS AS IDENTITY PRIMARY KEY,
  sport_id          smallint NOT NULL REFERENCES sports(id),
  season_id         bigint REFERENCES seasons(id),
  venue_id          bigint REFERENCES venues(id),
  round_label       text,
  starts_at         timestamptz NOT NULL,          -- always UTC on the wire
  status            text NOT NULL DEFAULT 'scheduled'
                    CHECK (status IN ('scheduled','live','paused','finished',
                                      'postponed','cancelled','abandoned','walkover')),
  status_detail     text,                          -- 'halftime', 'rain delay', 'retired'
  attributes        jsonb NOT NULL DEFAULT '{}',
  source_updated_at timestamptz,
  created_at        timestamptz NOT NULL DEFAULT now(),
  updated_at        timestamptz NOT NULL DEFAULT now()
);</p>
<p>CREATE INDEX fixtures_starts_idx      ON fixtures (sport_id, starts_at);
CREATE INDEX fixtures_season_idx      ON fixtures (season_id, starts_at);
CREATE INDEX fixtures_live_idx        ON fixtures (sport_id) WHERE status IN ('live','paused');</p>
<p>The partial index on live matches is a small trick with a large payoff: at any moment only a tiny fraction of your fixtures are live, so the index stays tiny while making the hottest query in your product cheap.</p>
<p>Participants attach through slots:</p>
<p>sql
CREATE TABLE fixture_competitors (
  fixture_id    bigint NOT NULL REFERENCES fixtures(id) ON DELETE CASCADE,
  slot          smallint NOT NULL,                 -- 1, 2, 3 ... stable position
  competitor_id bigint NOT NULL REFERENCES competitors(id),
  role          text,                              -- 'home', 'away', 'red_corner', 'runner', NULL
  entry_number  text,                              -- bib, draw, saddle cloth
  PRIMARY KEY (fixture_id, slot),
  UNIQUE (fixture_id, competitor_id)
);</p>
<p>CREATE INDEX fc_competitor_idx ON fixture_competitors (competitor_id, fixture_id);</p>
<p>Football uses slots 1 and 2 with roles home and away. A combat bout uses two slots with corner roles. Golf uses up to a hundred and fifty slots. A horse race uses a dozen. The query "all fixtures for this competitor" is identical in every case, and that is the point.</p>
<p>Do not infer home and away from slot order in application code. Read the role column, because some sports (neutral-venue finals, tennis, esports) have no meaningful home side.</p>
<p>Part 7: Scores, periods and final results</p>
<p>This is where sports diverge most, so give it the most careful design. The trick is to separate what happened per period from how it ended.</p>
<p>sql
CREATE TABLE fixture_periods (
  fixture_id  bigint NOT NULL REFERENCES fixtures(id) ON DELETE CASCADE,
  slot        smallint NOT NULL,
  period_kind text NOT NULL,             -- 'half','quarter','set','innings','round','map'
  period_no   smallint NOT NULL,
  score       integer,                   -- goals, points, games, runs, strokes
  details     jsonb NOT NULL DEFAULT '{}',   -- {"wickets":4,"overs":"19.2"} / {"tiebreak":7}
  PRIMARY KEY (fixture_id, slot, period_kind, period_no),
  FOREIGN KEY (fixture_id, slot) REFERENCES fixture_competitors (fixture_id, slot) ON DELETE CASCADE
);</p>
<p>CREATE TABLE fixture_results (
  fixture_id  bigint NOT NULL REFERENCES fixtures(id) ON DELETE CASCADE,
  slot        smallint NOT NULL,
  final_score integer,
  rank        smallint,                  -- 1 = winner / first place
  outcome     text CHECK (outcome IN ('win','loss','draw','no_result','dnf','dq','withdrawn')),
  method      text,                      -- 'ko', 'decision', 'penalties', 'retirement', 'dls'
  details     jsonb NOT NULL DEFAULT '{}',
  PRIMARY KEY (fixture_id, slot),
  FOREIGN KEY (fixture_id, slot) REFERENCES fixture_competitors (fixture_id, slot) ON DELETE CASCADE
);</p>
<p>Walk through how this absorbs the awkward sports:</p>
<p>Sport	How it maps
Football, Handball	Two slots, periods half, plus extra-time or penalty rows in period_kind
Basketball, American Football	Periods quarter, overtime as additional numbers
Ice Hockey	Periods period, shootout as a distinct kind, method records how it was decided
Tennis, Volleyball	Periods set, games in score, tiebreak points in details
Cricket	One period per innings, runs in score, wickets and overs in details, method = 'dls' where relevant
Baseball	Periods inning, hits and errors in details
Esports	Periods map or game, best-of format in fixture attributes
Combat Sports	Periods round, method holds KO, TKO, submission or decision
Golf	Periods round, score holds strokes, rank holds leaderboard position, outcome covers dnf and cuts
Horse Racing	One period or none, rank is finishing position, outcome covers non-finishers</p>
<p>Notice that "cricket needs overs and wickets" did not force a schema change. It went into details, because nothing in your product joins or filters on it. If you find yourself filtering on a JSONB key in a hot query, promote it to a real column. That is the whole rule.</p>
<p>Also notice the current score of a live match is not stored here as a special case. Live matches simply have period rows that are still changing, and fixtures.status = 'live'. One model, two moments in time.</p>
<p>Part 8: Events, the append-only heart of live data</p>
<p>Goals, cards, wickets, breaks of serve, knockdowns, birdies. Live products are built on a stream of events, so store them as one.</p>
<p>sql
CREATE TABLE fixture_events (
  id              bigint GENERATED ALWAYS AS IDENTITY PRIMARY KEY,
  fixture_id      bigint NOT NULL REFERENCES fixtures(id) ON DELETE CASCADE,
  seq             integer NOT NULL,                  -- order within the fixture
  kind            text NOT NULL,                     -- 'goal','card','wicket','point','birdie' ...
  slot            smallint,
  player_id       bigint REFERENCES players(id),
  clock           text,                              -- '67:12', '19.2', 'R3 2:41' (display only)
  occurred_at     timestamptz,
  payload         jsonb NOT NULL DEFAULT '{}',
  source_event_id text,
  voided_at       timestamptz,                       -- VAR overturned it, do not delete
  created_at      timestamptz NOT NULL DEFAULT now(),
  UNIQUE (fixture_id, seq),
  UNIQUE (fixture_id, source_event_id)
);</p>
<p>CREATE INDEX events_fixture_idx ON fixture_events (fixture_id, seq);</p>
<p>Three deliberate choices here:</p>
<p>voided_at instead of DELETE. A goal ruled out is still part of the record, and downstream clients that already displayed it need a correction, not silence.
clock is text. A football minute, a cricket over, and a boxing round-and-time are not the same type. Sorting uses seq, never the clock.
Two unique constraints give you idempotency. If the feed redelivers event 42, the second insert fails harmlessly or becomes a no-op.
Part 9: Provider IDs and the external mapping table</p>
<p>Now the rule from Part 1. Never let a provider's identifier become your key. Keep a translation table instead.</p>
<p>sql
CREATE TABLE external_ids (
  provider     text NOT NULL,           -- 'orbistats'
  entity_type  text NOT NULL,           -- 'fixture','competitor','player','competition',...
  external_id  text NOT NULL,
  internal_id  bigint NOT NULL,
  first_seen   timestamptz NOT NULL DEFAULT now(),
  PRIMARY KEY (provider, entity_type, external_id)
);</p>
<p>CREATE INDEX external_ids_internal_idx ON external_ids (entity_type, internal_id);</p>
<p>Why bother if you only use one provider? Because the day you add a second one, for odds or for a sport the first covers thinly, you will need to say "these two records are the same match". With this table that is two rows. Without it, it is a migration.</p>
<p>The lookup-or-create helper needs care, because two ingestion workers can see the same new fixture at once. Serialize on the key with a transaction-scoped advisory lock:</p>
<p>ts
// ids.ts
import type { PoolClient } from 'pg';</p>
<p>export async function resolveOrCreate(
  db: PoolClient,
  provider: string,
  entityType: string,
  externalId: string,
  create: () =&gt; Promise,
): Promise {
  await db.query('SELECT pg_advisory_xact_lock(hashtextextended($1, 0))', [
    <code>${provider}:${entityType}:${externalId}</code>,
  ]);</p>
<p>  const found = await db.query(
    <code>SELECT internal_id FROM external_ids      WHERE provider = $1 AND entity_type = $2 AND external_id = $3</code>,
    [provider, entityType, externalId],
  );
  if (found.rowCount) return Number(found.rows[0].internal_id);</p>
<p>  const internalId = await create();
  await db.query(
    <code>INSERT INTO external_ids (provider, entity_type, external_id, internal_id)      VALUES ($1, $2, $3, $4)</code>,
    [provider, entityType, externalId, internalId],
  );
  return internalId;
}</p>
<p>Call it inside a transaction. The advisory lock releases automatically on commit or rollback, so there is nothing to clean up.</p>
<p>Part 10: Statistics, the EAV versus JSONB debate</p>
<p>Statistics are where teams over-engineer. There are three common approaches:</p>
<p>One wide table per sport with a column per stat. Fast to query, miserable to evolve, and you will have thirteen of them.
One JSONB blob per fixture. Trivial to store, hard to aggregate, easy to corrupt.
A narrow key-value table with a definitions catalog. Slightly more rows, but one query shape for every sport.</p>
<p>For a startup covering many sports, I recommend option 3 with a defined catalog:</p>
<p>sql
CREATE TABLE stat_definitions (
  id        integer GENERATED ALWAYS AS IDENTITY PRIMARY KEY,
  sport_id  smallint NOT NULL REFERENCES sports(id),
  key       text NOT NULL,                   -- 'shots_on_target', 'aces', 'strokes_gained'
  name      text NOT NULL,
  scope     text NOT NULL CHECK (scope IN ('competitor','player')),
  unit      text,                            -- 'count','percent','seconds','metres'
  UNIQUE (sport_id, key, scope)
);</p>
<p>CREATE TABLE competitor_fixture_stats (
  fixture_id  bigint NOT NULL REFERENCES fixtures(id) ON DELETE CASCADE,
  slot        smallint NOT NULL,
  stat_id     integer NOT NULL REFERENCES stat_definitions(id),
  value       numeric NOT NULL,
  PRIMARY KEY (fixture_id, slot, stat_id),
  FOREIGN KEY (fixture_id, slot) REFERENCES fixture_competitors (fixture_id, slot) ON DELETE CASCADE
);</p>
<p>CREATE TABLE player_fixture_stats (
  fixture_id  bigint NOT NULL REFERENCES fixtures(id) ON DELETE CASCADE,
  player_id   bigint NOT NULL REFERENCES players(id),
  slot        smallint NOT NULL,
  stat_id     integer NOT NULL REFERENCES stat_definitions(id),
  value       numeric NOT NULL,
  PRIMARY KEY (fixture_id, player_id, stat_id)
);</p>
<p>CREATE INDEX pfs_player_idx ON player_fixture_stats (player_id, stat_id);</p>
<p>Two separate tables, not one with nullable subject columns, because a primary key cannot include a nullable column and you would end up with awkward sentinel values.</p>
<p>Aggregations become uniform across sports:</p>
<p>sql
-- Average of any stat for a player across a season
SELECT p.name, avg(s.value) AS avg_value
FROM player_fixture_stats s
JOIN stat_definitions d ON d.id = s.stat_id AND d.key = 'shots_on_target'
JOIN fixtures f ON f.id = s.fixture_id
JOIN players p ON p.id = s.player_id
WHERE f.season_id = $1
GROUP BY p.name
ORDER BY avg_value DESC
LIMIT 20;</p>
<p>The catalog is a feature, not a chore. It is also where you record a stat that a provider names inconsistently between sports, so your UI never learns about the mess. The Sports Statistics API describes team, player and match-level statistics with sport-specific metrics, which is exactly what this catalog absorbs.</p>
<p>Part 11: Standings as snapshots</p>
<p>A standings table that only stores "now" cannot answer "what was the table after round 12". Store it as a snapshot keyed by the round it describes.</p>
<p>sql
CREATE TABLE standings_rows (
  season_id     bigint NOT NULL REFERENCES seasons(id),
  group_key     text NOT NULL DEFAULT 'overall',   -- 'overall','group-a','east','home'
  as_of_round   integer NOT NULL,
  competitor_id bigint NOT NULL REFERENCES competitors(id),
  rank          smallint NOT NULL,
  played        smallint NOT NULL DEFAULT 0,
  points        numeric,
  details       jsonb NOT NULL DEFAULT '{}',       -- W/D/L, goal difference, NRR, form
  PRIMARY KEY (season_id, group_key, as_of_round, competitor_id)
);</p>
<p>Sport-specific tiebreakers (goal difference, net run rate, head-to-head) belong in details, while rank is whatever the provider computed. Do not recompute league tables yourself unless you are prepared to own every tiebreak rule of every competition.</p>
<p>Part 12: Odds without floating-point pain</p>
<p>Odds data has two lifetimes that you must not mix: what the price is right now, and every price it has ever been. Model them separately.</p>
<p>sql
CREATE TABLE bookmakers (
  id   integer GENERATED ALWAYS AS IDENTITY PRIMARY KEY,
  slug text NOT NULL UNIQUE,
  name text NOT NULL
);</p>
<p>CREATE TABLE market_types (
  id       integer GENERATED ALWAYS AS IDENTITY PRIMARY KEY,
  key      text NOT NULL UNIQUE,      -- 'moneyline','1x2','handicap','total','player_prop','outright'
  name     text NOT NULL
);</p>
<p>CREATE TABLE markets (
  id             bigint GENERATED ALWAYS AS IDENTITY PRIMARY KEY,
  fixture_id     bigint NOT NULL REFERENCES fixtures(id) ON DELETE CASCADE,
  market_type_id integer NOT NULL REFERENCES market_types(id),
  period         text NOT NULL DEFAULT 'full',     -- 'full','first_half','set_1'
  line           numeric,                          -- 2.5 goals, -1.5 handicap, NULL if none
  UNIQUE NULLS NOT DISTINCT (fixture_id, market_type_id, period, line)
);</p>
<p>CREATE TABLE selections (
  id            bigint GENERATED ALWAYS AS IDENTITY PRIMARY KEY,
  market_id     bigint NOT NULL REFERENCES markets(id) ON DELETE CASCADE,
  key           text NOT NULL,                     -- 'home','draw','away','over','under','yes'
  slot          smallint,                          -- link to a fixture competitor when relevant
  player_id     bigint REFERENCES players(id),
  UNIQUE (market_id, key, slot, player_id) 
);</p>
<p>UNIQUE NULLS NOT DISTINCT needs PostgreSQL 15 or newer. On older versions, use a unique index over coalesce(line, -999999). Note the UNIQUE on selections has nullable columns too, so apply the same treatment, or use an expression index there.</p>
<p>Now the two price tables:</p>
<p>sql
-- The latest price per selection per bookmaker: small, hot, upserted constantly
CREATE TABLE odds_current (
  selection_id bigint  NOT NULL REFERENCES selections(id) ON DELETE CASCADE,
  bookmaker_id integer NOT NULL REFERENCES bookmakers(id),
  price        numeric(9,3) NOT NULL CHECK (price &gt; 1),   -- decimal odds
  is_live      boolean NOT NULL DEFAULT false,
  suspended    boolean NOT NULL DEFAULT false,
  updated_at   timestamptz NOT NULL,
  PRIMARY KEY (selection_id, bookmaker_id)
);</p>
<p>-- Every change ever: append-only, partitioned by time so old data can be dropped cheaply
CREATE TABLE odds_ticks (
  selection_id bigint  NOT NULL,
  bookmaker_id integer NOT NULL,
  recorded_at  timestamptz NOT NULL,
  price        numeric(9,3) NOT NULL,
  suspended    boolean NOT NULL DEFAULT false,
  PRIMARY KEY (selection_id, bookmaker_id, recorded_at)
) PARTITION BY RANGE (recorded_at);</p>
<p>CREATE TABLE odds_ticks_2026_10 PARTITION OF odds_ticks
  FOR VALUES FROM ('2026-10-01') TO ('2026-11-01');</p>
<p>Rules for odds that save real money:</p>
<p>Store decimal odds as numeric, never float. Convert American and fractional formats at the edge, and keep one canonical format internally.
Compute implied probability in queries or views, not in storage. It is 1 / price, and margin (overround) is the sum across a market's selections minus one.
Keep the suspended flag. A market that is closed is different from a market with no price.
Automate partition creation and retention. Dropping a month-old partition is instant, while deleting a hundred million rows is an incident.
Do not promise depth you do not have. The Odds API describes normalized pre-match and live markets, and market depth can vary by competition. Your market_types seed should grow as you verify what each competition really provides.
Part 13: Idempotent, ordered ingestion</p>
<p>Feeds lie in three predictable ways: they repeat, they reorder, and they correct themselves. The write path must survive all three. Guard updates with the source timestamp so an older message can never overwrite a newer one:</p>
<p>sql
UPDATE fixtures
SET status            = $2,
    status_detail     = $3,
    source_updated_at = $4,
    updated_at        = now()
WHERE id = $1
  AND (source_updated_at IS NULL OR source_updated_at &lt;= $4);</p>
<p>And for periods and current odds, use upserts with the same guard:</p>
<p>sql
INSERT INTO fixture_periods (fixture_id, slot, period_kind, period_no, score, details)
VALUES ($1, $2, $3, $4, $5, $6)
ON CONFLICT (fixture_id, slot, period_kind, period_no)
DO UPDATE SET score = EXCLUDED.score, details = EXCLUDED.details;</p>
<p>INSERT INTO odds_current (selection_id, bookmaker_id, price, is_live, suspended, updated_at)
VALUES ($1, $2, $3, $4, $5, $6)
ON CONFLICT (selection_id, bookmaker_id)
DO UPDATE SET price = EXCLUDED.price, is_live = EXCLUDED.is_live,
              suspended = EXCLUDED.suspended, updated_at = EXCLUDED.updated_at
WHERE odds_current.updated_at &lt;= EXCLUDED.updated_at;</p>
<p>Keep raw input for replay, with a hash to skip exact duplicates:</p>
<p>sql
CREATE TABLE raw_ingest (
  id          bigint GENERATED ALWAYS AS IDENTITY PRIMARY KEY,
  provider    text NOT NULL,
  resource    text NOT NULL,                  -- 'fixtures','live','odds'
  received_at timestamptz NOT NULL DEFAULT now(),
  payload_hash text NOT NULL,
  payload     jsonb NOT NULL
);
CREATE INDEX raw_ingest_lookup ON raw_ingest (provider, resource, payload_hash);</p>
<p>Retain it for days or weeks, not forever. It is a debugging tool, not an archive.</p>
<p>Part 14: The adapter, mapping payloads into the schema</p>
<p>Validate at the boundary, and translate vendor vocabulary into your own exactly once. Here is a status mapper and a validated fixture shape. Every field name is a placeholder, so check the sandbox.</p>
<p>ts
// adapter/fixture.ts
import { z } from 'zod';</p>
<p>const RawFixture = z.object({
  id: z.union([z.string(), z.number()]).transform(String),
  sport: z.string(),
  start_time: z.string(),                       // ISO 8601, assumed UTC
  status: z.string(),
  competitors: z.array(z.object({
    id: z.union([z.string(), z.number()]).transform(String),
    name: z.string(),
    role: z.string().optional(),
  })).min(1),
  updated_at: z.string().optional(),
});</p>
<p>export type NormalizedStatus =
  'scheduled' | 'live' | 'paused' | 'finished' |
  'postponed' | 'cancelled' | 'abandoned' | 'walkover';</p>
<p>const STATUS_MAP: Record&lt;string, NormalizedStatus&gt; = {
  scheduled: 'scheduled', not_started: 'scheduled',
  live: 'live', in_play: 'live',
  halftime: 'paused', break: 'paused', delayed: 'paused',
  finished: 'finished', ended: 'finished', final: 'finished',
  postponed: 'postponed', cancelled: 'cancelled', canceled: 'cancelled',
  abandoned: 'abandoned', walkover: 'walkover',
};</p>
<p>export function toStatus(raw: string): NormalizedStatus {
  const s = STATUS_MAP[raw.toLowerCase()];
  if (!s) {
    // Unknown vocabulary must be loud, never silently coerced
    throw new Error(<code>unmapped status: ${raw}</code>);
  }
  return s;
}</p>
<p>export function parseFixture(input: unknown) {
  const r = RawFixture.parse(input);
  return {
    externalId: r.id,
    sportSlug: r.sport,
    startsAt: new Date(r.start_time),
    status: toStatus(r.status),
    competitors: r.competitors.map((c, i) =&gt; ({
      slot: i + 1,
      externalId: c.id,
      name: c.name,
      role: c.role ?? null,
    })),
    sourceUpdatedAt: r.updated_at ? new Date(r.updated_at) : new Date(),
  };
}</p>
<p>The unknown-status behaviour is deliberate. A vendor that adds a status called interrupted should trigger an alert and a mapping update, not a database full of fixtures with a bogus state. Read the changelog as part of your release routine, and write contract tests against saved sandbox responses so a changed field name fails your CI instead of your users.</p>
<p>Part 15: Time, the source of quiet bugs
Store timestamptz and treat everything on the wire as UTC.
Store the venue's IANA timezone separately, and convert only at the display edge.
Never compute "match day" by truncating a UTC timestamp, because a late kickoff in India is a different UTC date than it is locally. Derive the local date from the venue timezone when you need it.
Track rescheduling. When starts_at changes, write the old value to an audit table, because odds, notifications and user reminders all depend on the original.
sql
CREATE TABLE fixture_schedule_changes (
  fixture_id  bigint NOT NULL REFERENCES fixtures(id) ON DELETE CASCADE,
  changed_at  timestamptz NOT NULL DEFAULT now(),
  old_start   timestamptz NOT NULL,
  new_start   timestamptz NOT NULL,
  reason      text
);
Part 16: Corrections, voids and history</p>
<p>Sports data is revised constantly: a goal is reassigned, a stat is corrected, a result is amended after a protest. Your schema should make corrections cheap and history recoverable.</p>
<p>Events are voided, not deleted (Part 8).
Final results are overwritten by the upsert guard, but keep the previous value in an audit table if your product shows "corrected" badges or feeds betting settlement.
For long-term analysis, backfill archives from the Historical Sports Data API into the same tables rather than a separate archive schema. One model means one set of queries for both.
Part 17: What to build first (the startup version)</p>
<p>You do not need every table on day one. A sensible order:</p>
<p>Stage 1, ship a scores app for one sport (about nine tables): sports, competitions, seasons, competitors, fixtures, fixture_competitors, fixture_periods, fixture_results, external_ids.</p>
<p>Stage 2, make it live: fixture_events, raw_ingest, the partial live index, and idempotent upserts. Wire it to the WebSocket API or Webhooks instead of polling.</p>
<p>Stage 3, make it analytical: players, competitor_members, stat_definitions, the two stats tables, standings_rows.</p>
<p>Stage 4, make it commercial: the odds tables, partitioned ticks, and a second provider if you need one.</p>
<p>Stage 5, go multi-sport: add sports as rows, then extend period_kind values, stat_definitions and event kind values per sport. If you followed the slot model, this stage is mostly data entry.</p>
<p>A quick way to feel this in practice is to build the smallest useful client first, for example the 50-line live scoreboard walkthrough, then design tables around the payloads you actually saw. If your backend is not Node, there are also Python (FastAPI) and .NET Core examples.</p>
<p>Part 18: Ten schema mistakes that hurt later
Using provider IDs as primary keys.
Hard-coding home_team_id and away_team_id.
Storing scores only on the fixture row, so periods and innings have nowhere to live.
Using float for odds, then wondering why probabilities do not sum.
One giant JSONB column for everything, and then filtering on it in hot queries.
One table per sport, which multiplies every future change by thirteen.
Deleting corrected events, so clients can never be told what changed.
Deriving "finished" from starts_at + 3 hours.
Storing local times, or trusting a date column without a timezone.
Skipping the raw payload store, then being unable to replay after a mapping bug.
Part 19: Testing your schema against reality
Load the sandbox. Save real responses from the sandbox as fixtures for automated tests, and run the full mapper against every supported sport.
Replay a match day. Feed a recorded sequence, including duplicates and out-of-order messages, and assert the final rows are identical to an ordered run.
Property-test idempotency. Applying any update twice must equal applying it once.
Check the awkward sports first. If tennis doubles, a cricket rain-shortened match, a golf tournament with a cut, and a horse race with a non-runner all fit without special cases, the model is sound.
Verify coverage per sport. The website shows 13 sports, but individual product pages can differ, so check pricing and plan details and the documentation before promising a feature for a specific sport to your own users.
Part 20: A note on licensing</p>
<p>If your product stores upstream data and shows it to end users, read the provider's terms on caching, storage duration, and redistribution before you design retention. Some agreements limit how long data may be kept or whether raw data may be re-served through your own API. Orbistats positions enterprise plans around custom feeds, dedicated infrastructure and SLA requirements, so raise questions early. The developer hub and the About page are useful starting points for understanding what the provider offers and how it positions itself.</p>
<p>A repeatable checklist
Seed sports as rows, one per supported sport.
Model competition as fixtures, slot-based competitors, periods, results and events.
Keep internal IDs separate from provider IDs through external_ids.
Put filterable fields in columns, and display-only oddities in JSONB.
Use a narrow stats table with a definitions catalog.
Store standings as round snapshots.
Split odds into odds_current and partitioned odds_ticks, using numeric decimal prices.
Make every write idempotent, and guard updates with the source timestamp.
Validate at the edge, and fail loudly on unknown vocabulary.
Keep raw input briefly, replay it when things break, and test with real sandbox payloads.
Final thoughts</p>
<p>A good sports schema is not clever. It is a small set of ideas applied consistently: competitors in slots, periods that carry scores, events that append, prices that tick, and identifiers that belong to you. Build the boring core once, and the thirteenth sport costs you an afternoon instead of a quarter. Start with one sport, keep the boundaries clean, and let real payloads from the sandbox correct your assumptions early.</p>
<p>If you have modelled a sport that broke your schema, tell me which one in the comments. Cricket and golf are usually the first to find the flaws.</p>
]]></content:encoded></item><item><title><![CDATA[How We Cut Sports API Latency by 80%: A Case Study]]></title><description><![CDATA[Most latency posts start with a trick. This one starts with a stopwatch, because in a live sports product the fastest way to waste a week is to optimize the wrong layer. The lesson I care about most i]]></description><link>https://bettechmagnetics.hashnode.dev/how-we-cut-sports-api-latency-by-80-a-case-study</link><guid isPermaLink="true">https://bettechmagnetics.hashnode.dev/how-we-cut-sports-api-latency-by-80-a-case-study</guid><dc:creator><![CDATA[Vijay Choudhary]]></dc:creator><pubDate>Tue, 29 Sep 2026 13:21:36 GMT</pubDate><enclosure url="https://cdn.hashnode.com/uploads/covers/6a47ac8dec0b8a17b4dc915d/b4316089-cdd6-44c0-8b22-dc053e868587.png" length="0" type="image/jpeg"/><content:encoded><![CDATA[<p>Most latency posts start with a trick. This one starts with a stopwatch, because in a live sports product the fastest way to waste a week is to optimize the wrong layer. The lesson I care about most is simple: measure the path you own, fix the biggest slice first, and re-measure after every change.</p>
<p>About the numbers in this post</p>
<p>The system below is a realistic reference build, not a claim about any specific company's traffic. The figures in the results table are illustrative. Replace them with your own measurements before you publish or quote them anywhere. The code, the measurement method and the order of fixes are the reusable parts.</p>
<p>I use Orbistats as the upstream provider because its public docs describe REST, WebSocket and Webhooks delivery, and the same patterns apply to any licensed feed. The vendor's sub-50ms feed claim is the vendor's own statement. It is not something I measured, and you should not treat it as a benchmark of your end-to-end path.</p>
<p>The situation</p>
<p>Imagine a live scores product with a web app, a mobile app and a few embedded widgets. It shows fixtures, live scores, match events and statistics across several sports. Orbistats currently advertises 13 sports on its homepage, each with a dedicated page such as tennis, and you should check the documentation for the exact list before promising any sport to your users.</p>
<p>The first version was the obvious one: every client polled our backend every two seconds, and our backend called the upstream REST API on every request.</p>
<p>ts
// v1: the naive version (do not ship this)
app.get('/api/live/:sport', async (req, res) =&gt; {
  const r = await fetch(<code>https://api.orbistats.com/v1/${req.params.sport}/live</code>, {
    headers: { Authorization: <code>Bearer ${process.env.UPSTREAM_KEY}</code> },
  });
  res.status(r.status).json(await r.json());
});</p>
<p>Symptoms:</p>
<p>Slow tail: p95 of the live endpoint was far worse than the median, especially on match days.
Upstream pressure: upstream request volume grew linearly with the number of open browser tabs.
Stale-feeling scores: a goal could appear up to one polling interval late, on top of request latency.
Fragile behaviour: one slow upstream response made every waiting client slow.
Step 0: Define what "latency" means</p>
<p>"Latency" is at least three different numbers, and mixing them is how teams fool themselves:</p>
<p>Request latency: how long a client waits for a REST response.
Update latency (freshness): how long after something happens in the match a user sees it.
Fan-out lag: how long an event sits inside your system between arriving from upstream and being written to a client socket.</p>
<p>You can measure the first and third precisely, because both clocks are yours. The second includes the provider's delay, which you can only bound, not observe directly. Keep them separate in every chart.</p>
<p>Step 1: Measure before touching anything
Break one request into phases</p>
<p>Start with curl. It shows whether time goes to DNS, TCP, TLS, waiting for the server, or downloading.</p>
<p>bash
curl -s -o /dev/null <br />  -H "Authorization: Bearer $UPSTREAM_KEY" <br />  -w "dns:%{time_namelookup}s connect:%{time_connect}s tls:%{time_appconnect}s ttfb:%{time_starttransfer}s total:%{time_total}s size:%{size_download}B\n" <br />  <a href="https://api.orbistats.com/v1/football/live">https://api.orbistats.com/v1/football/live</a></p>
<p>Run it 20 times and look at the spread, not one lucky run. If connect plus TLS is a large share of total, connection reuse will help. If ttfb dominates, the wait is upstream or in your own processing.</p>
<p>Put a histogram on your own endpoints
ts
// metrics.ts
import client from 'prom-client';</p>
<p>export const httpLatency = new client.Histogram({
  name: 'api_request_seconds',
  help: 'REST latency by route and cache status',
  labelNames: ['route', 'cache', 'status'],
  buckets: [0.005, 0.01, 0.025, 0.05, 0.1, 0.25, 0.5, 1, 2.5],
});</p>
<p>export const fanoutLag = new client.Histogram({
  name: 'stream_fanout_lag_seconds',
  help: 'Time from upstream arrival to client socket write',
  labelNames: ['sport'],
  buckets: [0.001, 0.005, 0.01, 0.025, 0.05, 0.1, 0.25, 0.5],
});
ts
// timing middleware
app.use((req, res, next) =&gt; {
  const end = httpLatency.startTimer();
  res.on('finish', () =&gt; {
    end({
      route: req.route?.path ?? 'unknown',
      cache: String(res.getHeader('X-Cache') ?? 'na'),
      status: String(res.statusCode),
    });
  });
  next();
});</p>
<p>Always label by cache status. Without that label, hits and misses blur into one average that describes nothing.</p>
<p>Generate load that looks like match day
bash
npx autocannon -c 200 -d 30 -H "Authorization: Bearer $TENANT_KEY" <br />  <a href="http://localhost:3000/api/live/football">http://localhost:3000/api/live/football</a></p>
<p>Record p50, p95 and p99. The median hides the pain, and users feel the tail.</p>
<p>Step 2: Stop making each client cost an upstream call</p>
<p>The biggest structural win is usually not a micro-optimization. It is decoupling how many users you have from how many upstream calls you make.</p>
<p>Fix A: One shared upstream stream instead of per-client polling</p>
<p>Orbistats documents a persistent WebSocket API and a Webhooks API so you do not have to poll. Open one upstream connection per service instance and fan out internally. Copy the real subscribe payload and host from the docs, and read the changelog because streaming details can change.</p>
<p>ts
// upstream.ts
import WebSocket from 'ws';</p>
<p>type Handler = (raw: any, receivedAt: number) =&gt; void;
const handlers: Handler[] = [];
export const onUpstream = (h: Handler) =&gt; handlers.push(h);</p>
<p>let attempt = 0;</p>
<p>export function connectUpstream() {
  const ws = new WebSocket(process.env.UPSTREAM_WS_URL!, {
    headers: { Authorization: <code>Bearer ${process.env.UPSTREAM_KEY}</code> },
    perMessageDeflate: false, // small frames: compression costs more CPU than it saves
  });</p>
<p>  let alive = true;
  let hb: NodeJS.Timeout;</p>
<p>  ws.on('open', () =&gt; {
    attempt = 0;
    // PLACEHOLDER: use the exact subscribe message from the WebSocket docs
    ws.send(JSON.stringify({ action: 'subscribe', channel: 'live' }));
    hb = setInterval(() =&gt; {
      if (!alive) return ws.terminate();
      alive = false;
      ws.ping();
    }, 30_000);
  });</p>
<p>  ws.on('pong', () =&gt; { alive = true; });</p>
<p>  ws.on('message', buf =&gt; {
    const receivedAt = Date.now();          // stamp immediately, before any parsing
    let msg: any;
    try { msg = JSON.parse(buf.toString()); } catch { return; }
    const items = Array.isArray(msg.data) ? msg.data : [msg.data ?? msg];
    for (const raw of items) handlers.forEach(h =&gt; h(raw, receivedAt));
  });</p>
<p>  ws.on('close', () =&gt; {
    clearInterval(hb);
    const delay = Math.min(30_000, 1000 &lt;&lt; attempt++) + Math.floor(Math.random() * 500);
    setTimeout(connectUpstream, delay);
  });
  ws.on('error', e =&gt; console.error('upstream error', e.message));
}</p>
<p>Two details matter here. The timestamp is taken first, so fan-out lag includes your own parse time. And reconnects use exponential backoff with jitter, so a network blip does not become a reconnect storm.</p>
<p>Fix B: Keep a live state store, and serve REST from it</p>
<p>Once events flow in, keep the latest state per match in memory. Your "live" REST endpoint stops being a proxy and becomes a memory read.</p>
<p>ts
// state.ts
export interface MatchState {
  matchId: string;
  sport: string;
  data: Record&lt;string, unknown&gt;;
  updatedAt: number;
}</p>
<p>const live = new Map&lt;string, MatchState&gt;();</p>
<p>export function applyUpdate(sport: string, raw: any) {
  const matchId = String(raw.match_id ?? raw.fixture_id ?? raw.id ?? '');
  if (!matchId) return null;
  const prev = live.get(matchId);
  const next: MatchState = {
    matchId,
    sport,
    data: { ...(prev?.data ?? {}), ...raw },
    updatedAt: Date.now(),
  };
  live.set(matchId, next);
  return { prev, next };
}</p>
<p>export const liveBySport = (sport: string) =&gt;
  [...live.values()].filter(m =&gt; m.sport === sport);
ts
app.get('/api/live/:sport', (req, res) =&gt; {
  res.setHeader('X-Cache', 'state');
  res.json(liveBySport(req.params.sport));   // no network hop, no upstream call
});</p>
<p>Field names above (match_id, fixture_id) are assumptions. Compare a real payload from the public API sandbox and adjust only the adapter. Also seed the store on startup with a REST snapshot from the Live Scores API, so a fresh instance is not empty until the next event arrives.</p>
<p>Fix C: Coalesce and use stale-while-revalidate for everything that is not live</p>
<p>Fixtures, standings and teams from the Sports Data API change slowly. Collapse identical concurrent requests into one, and serve slightly stale data while refreshing in the background.</p>
<p>ts
// cache.ts
const inflight = new Map&lt;string, Promise&gt;();
const store = new Map&lt;string, { v: unknown; fresh: number; stale: number }&gt;();</p>
<p>function coalesce(key: string, load: () =&gt; Promise): Promise {
  const p = inflight.get(key);
  if (p) return p as Promise;
  const n = load().finally(() =&gt; inflight.delete(key));
  inflight.set(key, n);
  return n;
}</p>
<p>export async function swr(key: string, freshMs: number, staleMs: number, load: () =&gt; Promise) {
  const now = Date.now();
  const hit = store.get(key);
  if (hit &amp;&amp; now &lt; hit.fresh) return { value: hit.v as T, cache: 'hit' };</p>
<p>  const refresh = () =&gt; coalesce(key, load).then(v =&gt; {
    const t = Date.now();
    store.set(key, { v, fresh: t + freshMs, stale: t + freshMs + staleMs });
    return v;
  });</p>
<p>  if (hit &amp;&amp; now &lt; hit.stale) {
    refresh().catch(() =&gt; {});
    return { value: hit.v as T, cache: 'stale' };
  }
  return { value: await refresh(), cache: 'miss' };
}</p>
<p>Pick windows per data type: seconds for live-adjacent data, minutes for standings, hours for team and player metadata, and far longer for archives from the Historical Sports Data API. If you run more than one node, move this store into Redis so nodes share hits.</p>
<p>Step 3: Remove the connection tax</p>
<p>Every new upstream request over a fresh connection pays for DNS, TCP and TLS. Node's built-in fetch reuses connections, but explicit tuning makes behaviour predictable under load.</p>
<p>ts
// http.ts
import { Agent, setGlobalDispatcher } from 'undici';</p>
<p>setGlobalDispatcher(new Agent({
  keepAliveTimeout: 30_000,
  keepAliveMaxTimeout: 60_000,
  connections: 32,        // per origin; size to your concurrency, not to a guess
  connect: { timeout: 3_000 },
}));</p>
<p>export const upstreamFetch = (url: string) =&gt;
  fetch(url, {
    headers: { Authorization: <code>Bearer ${process.env.UPSTREAM_KEY}</code> },
    signal: AbortSignal.timeout(5_000),   // never wait forever on one upstream call
  });</p>
<p>Also give every upstream call a hard timeout. A missing timeout turns one slow response into a growing queue, and that queue is what wrecks p99.</p>
<p>Placement matters too. Run the service in a region close to the upstream API and close to most of your users, then re-run the curl phase breakdown. Moving compute is sometimes worth more than any code change, but only your own measurements can tell you which.</p>
<p>Step 4: Send fewer bytes</p>
<p>After the structural fixes, payload size becomes visible, especially on mobile networks.</p>
<p>Send deltas over the socket</p>
<p>A live score update usually changes one or two fields, yet full match objects are much larger. Send only what changed.</p>
<p>ts
// delta.ts
export function diff(prev: Record&lt;string, unknown&gt; | undefined, next: Record&lt;string, unknown&gt;) {
  const out: Record&lt;string, unknown&gt; = {};
  for (const k of Object.keys(next)) {
    if (JSON.stringify(prev?.[k]) !== JSON.stringify(next[k])) out[k] = next[k];
  }
  return out;
}</p>
<p>Send the full state once on subscribe, then deltas. Include a sequence number per match, so a client that detects a gap can request a fresh snapshot instead of silently drifting.</p>
<p>Add ETags and compression to REST
ts
import compression from 'compression';
import { createHash } from 'node:crypto';</p>
<p>app.use(compression({ threshold: 1024 }));</p>
<p>app.get('/api/fixtures/:sport', async (req, res) =&gt; {
  const { value, cache } = await swr(<code>fx:${req.params.sport}</code>, 30_000, 300_000,
    () =&gt; upstreamFetch(<code>https://api.orbistats.com/v1/${req.params.sport}/fixtures</code>).then(r =&gt; r.json()));</p>
<p>  const body = JSON.stringify(value);
  const etag = '"' + createHash('sha1').update(body).digest('hex') + '"';
  res.setHeader('ETag', etag);
  res.setHeader('X-Cache', cache);
  res.setHeader('Cache-Control', 'public, max-age=5, stale-while-revalidate=30');</p>
<p>  if (req.headers['if-none-match'] === etag) return res.status(304).end();
  res.type('json').send(body);
});</p>
<p>A 304 response has no body, so repeat visitors pay almost nothing. Only add public caching for data that is identical for every tenant. Anything entitlement-specific must not be shared through a public cache.</p>
<p>Step 5: Make fan-out cheap</p>
<p>With one upstream stream and many downstream sockets, the fan-out loop is your hot path. Two rules keep it fast: serialize once per event, not once per client, and never let a slow client block a fast one.</p>
<p>ts
// fanout.ts
import { WebSocket } from 'ws';
import { fanoutLag } from './metrics';</p>
<p>interface Client { ws: WebSocket; sports: Set; drops: number; }
export const clients = new Set();</p>
<p>const MAX_BUFFERED = 1_000_000;  // bytes queued on one socket before we treat it as slow</p>
<p>export function publish(sport: string, matchId: string, delta: object, receivedAt: number) {
  const frame = JSON.stringify({ sport, matchId, delta });   // once, shared by every client
  const stop = fanoutLag.startTimer({ sport });</p>
<p>  for (const c of clients) {
    if (!c.sports.has(sport) || c.ws.readyState !== WebSocket.OPEN) continue;
    if (c.ws.bufferedAmount &gt; MAX_BUFFERED) {
      if (++c.drops &gt; 200) c.ws.close(1013, 'slow_consumer');
      continue;
    }
    c.ws.send(frame);
  }
  stop();
  // Alternatively, observe the real elapsed time from the upstream timestamp:
  fanoutLag.observe({ sport }, (Date.now() - receivedAt) / 1000);
}</p>
<p>Keep the lag metric honest. It measures upstream arrival to your socket write, which is a number you own. It does not include the provider's own delay or the client's network.</p>
<p>For very bursty moments (a busy football Saturday), you can micro-batch frames in a 20 to 50 ms window to cut syscalls. That deliberately adds latency, so only do it if measurements show CPU is your bottleneck rather than time.</p>
<p>Step 6: Re-measure and read the result honestly</p>
<p>Re-run the same load test, on the same machine class, with the same duration. Change one thing at a time so you know what earned the improvement.</p>
<p>Metric (illustrative values, replace with yours)	Before	After	Change
p50 live endpoint	310 ms	8 ms	large drop
p95 live endpoint	820 ms	160 ms	about 80% lower
p99 live endpoint	2,400 ms	420 ms	about 82% lower
Upstream calls per minute at peak	18,000	40	decoupled from user count
Fan-out lag p95	not measured	12 ms	new visibility
Bytes per live update	6.4 KB	0.3 KB	deltas</p>
<p>Read this table the right way:</p>
<p>The 80% figure is a p95 request-latency result, not a claim that scores reach users 80% faster. Update freshness depends on the provider's delay too.
Which fix earned what is the real insight. In a build like this, removing per-client polling and serving from a live state store usually dominates. Connection tuning and compression are smaller, but they help the tail.
A p50 that improves more than p95 often means a cache is helping hits while misses still hurt. Check your X-Cache labels.
What did not work (and what we would watch)
Longer cache times on live data. They made graphs look great and scores feel wrong. Freshness is a product requirement, not just a metric.
Compression on tiny WebSocket frames. It cost CPU without saving meaningful bytes, so we turned it off for the stream.
Trusting one benchmark run. Run-to-run variance was large enough to fake a 20% "win". Repeat runs and compare distributions.
Ignoring upstream limits. Aggressive parallel warming can hit provider rate limits. Read the plan details on the pricing page and design for them.
Failure modes to design for
Failure	Effect	Response
Upstream socket drops	Live data goes stale	Reconnect with jitter, mark feed unhealthy, serve last known state with a stale flag
Upstream 5xx or timeout	REST errors	Stale-while-revalidate, hard timeouts, coalescing
Slow client	Memory growth	bufferedAmount check, drop, then disconnect
Cold start	Empty live store	Seed from a REST snapshot before accepting traffic
Provider changes a field	Broken parsing	Keep parsing in one adapter, and contract-test against the sandbox
Licensing note before you copy this design</p>
<p>If you re-serve upstream data to your own customers, check your provider agreement first. Many plans allow use inside your own app but not raw redistribution. Orbistats positions its enterprise plans around custom feeds, dedicated infrastructure and SLA requirements, so ask early. The About page and the developer hub are good starting points for understanding how it positions itself.</p>
<p>A repeatable checklist
Split latency into request latency, freshness and fan-out lag.
Phase-break one request with curl, and label histograms by cache status.
Replace per-client polling with one shared upstream stream.
Serve live REST from an in-memory state store seeded by a snapshot.
Coalesce and use stale-while-revalidate for slow-changing data.
Reuse connections and set hard timeouts on every upstream call.
Send deltas, add ETags, and compress only where it pays.
Serialize once per event and guard against slow consumers.
Re-measure with the same test, one change at a time.
Publish only numbers you measured yourself.
Where to go next
Build the smallest working client first with the 50-line live scoreboard walkthrough.
If your backend is Python or .NET, see the FastAPI wrapper example and the .NET Core client.
Add betting-style features with the Odds API, and give it its own cache windows because odds move faster than fixtures.
Add team analytics with the Sports Statistics API.
Explore real response shapes in the sandbox before you write a single parser.
Final thoughts</p>
<p>Latency work feels like a hunt for clever tricks, but the gains in this case came from boring structure: fewer upstream calls, a live state store, reused connections, smaller payloads and a fan-out loop that never waits on a slow client. Measure first, fix the largest slice, and then measure again. If your number improves, you can say so with a straight face.</p>
<p>If you have a latency win (or a failed experiment) from your own live-data system, share it in the comments. Those stories teach more than any benchmark.</p>
]]></content:encoded></item><item><title><![CDATA[Designing a Multi-Tenant Sports Data Platform: Lessons from Production]]></title><description><![CDATA[Most sports data articles show you how to call an API. This one is about the other side: what you build when you become the API for your own customers.
Imagine you run a product with many tenants. Som]]></description><link>https://bettechmagnetics.hashnode.dev/designing-a-multi-tenant-sports-data-platform-lessons-from-production</link><guid isPermaLink="true">https://bettechmagnetics.hashnode.dev/designing-a-multi-tenant-sports-data-platform-lessons-from-production</guid><category><![CDATA[architecture]]></category><category><![CDATA[System Design]]></category><category><![CDATA[websockets]]></category><category><![CDATA[api]]></category><category><![CDATA[Node.js]]></category><dc:creator><![CDATA[Vijay Choudhary]]></dc:creator><pubDate>Tue, 29 Sep 2026 13:13:20 GMT</pubDate><enclosure url="https://cdn.hashnode.com/uploads/covers/6a47ac8dec0b8a17b4dc915d/759f085d-6878-41b8-9428-06b74b1cea92.png" length="0" type="image/jpeg"/><content:encoded><![CDATA[<p>Most sports data articles show you how to call an API. This one is about the other side: what you build when you become the API for your own customers.</p>
<p>Imagine you run a product with many tenants. Some are hobby developers on a free plan, some are media sites that need widgets, some are analytics teams that pull history, and a few are demanding enterprise customers. All of them want live scores, fixtures, statistics and odds across many sports. You cannot open a separate upstream connection for each of them, you cannot let one heavy customer starve the rest, and you absolutely cannot show tenant A the data that only tenant B paid for.</p>
<p>This post walks through a practical design for that platform, with code for the parts that are easy to get wrong: authentication, rate limiting, caching, real-time fan-out, tenant isolation, and outbound webhooks.</p>
<p>A note on honesty before we start. This is a design walkthrough, not a benchmark report. Any numbers in the code (rates, quotas, cache times) are illustrative defaults you should tune against your own traffic. I use Orbistats as the example upstream provider because its public docs describe REST, WebSocket and Webhooks delivery, but the architecture applies to any licensed sports feed.</p>
<p>What multi-tenant means in this context</p>
<p>In a multi-tenant sports platform, one deployment serves many customers, and each customer (a tenant) gets its own:</p>
<ul>
<li>Identity: API keys that map to exactly one tenant</li>
<li>Limits: request rate, daily quota, and connection count</li>
<li>Entitlements: which sports, endpoints and real-time features they can use</li>
<li>Isolation: their configuration, usage and webhook secrets are invisible to others</li>
<li>Observability: you can see usage, errors and latency per tenant</li>
</ul>
<p>The hard part is doing this while sharing the expensive resource, which is the upstream data feed.</p>
<p>Requirements we will design for</p>
<ul>
<li>Serve REST and real-time streams to many tenants from one upstream subscription</li>
<li>Enforce per-tenant rate limits and daily quotas</li>
<li>Gate access by plan: sports, features, real-time or delayed</li>
<li>Never let one tenant's traffic degrade another's</li>
<li>Survive upstream outages and bad payloads gracefully</li>
<li>Support outbound webhooks with signatures and retries</li>
<li>Keep every provider-specific detail in one replaceable layer</li>
</ul>
<p>Architecture overview</p>
<pre><code>                        +--------------------+
  Upstream provider --&gt; | Ingestion layer    |
  (REST + WebSocket)    | adapter, normalize |
                        +---------+----------+
                                  |
                        +---------v----------+
                        | Event bus + cache  |
                        | (Redis or similar) |
                        +----+---------+-----+
                             |         |
                   +---------v--+   +--v-----------------+
                   | REST API   |   | Stream gateway     |
                   | auth, rate |   | per-tenant filter, |
                   | limit, SWR |   | slow consumer guard|
                   +-----+------+   +---------+----------+
                         |                    |
                   +-----v--------------------v-----+
                   | Tenants: apps, dashboards, bots |
                   +---------------------------------+
                             ^
                   +---------+----------+
                   | Webhook dispatcher |
                   | signed, retried    |
                   +--------------------+
</code></pre>
<p>The key idea: there is one ingestion layer that talks to the provider, and everything downstream reads from your own cache and bus. Tenants never touch the provider directly.</p>
<p>Layer 1: The ingestion layer and the provider adapter</p>
<p>Start with the layer that will change most often. The upstream provider offers several products, and it is worth mapping them to your internal needs. Orbistats describes a <a href="https://orbistats.com/api/sports-data-api.html">Sports Data API</a> for fixtures, results, standings, teams and players, a <a href="https://orbistats.com/api/live-scores-api.html">Live Scores API</a> for match state and events, a <a href="https://orbistats.com/api/sports-statistics-api.html">Sports Statistics API</a>, an <a href="https://orbistats.com/api/odds-api.html">Odds API</a> for normalized pre-match and live odds, and a <a href="https://orbistats.com/api/historical-sports-data-api.html">Historical Sports Data API</a> for research and backtesting. For push delivery it documents a persistent <a href="https://orbistats.com/api/websocket-api.html">WebSocket API</a> and the <a href="https://orbistats.com/api/webhooks-api.html">Webhooks API</a>.</p>
<p>Define your own interface so the rest of the platform never depends on a vendor shape:</p>
<pre><code class="language-ts">// provider.ts
export type Sport =
  | 'football' | 'basketball' | 'american-football' | 'cricket' | 'tennis'
  | 'baseball' | 'esports' | 'combat-sports' | 'volleyball' | 'handball'
  | 'ice-hockey' | 'golf' | 'horse-racing';

export interface MatchEvent {
  eventId: string;
  matchId: string;
  sport: Sport;
  kind: 'score' | 'status' | 'odds' | 'stat';
  payload: Record&lt;string, unknown&gt;;
  receivedAt: number;
}

export interface SportsProvider {
  snapshot(sport: Sport, resource: string, params: Record&lt;string, string&gt;): Promise&lt;unknown&gt;;
  onEvent(handler: (e: MatchEvent) =&gt; void): void;
  start(): void;
}
</code></pre>
<p>Orbistats lists 13 sports on its <a href="https://orbistats.com/">homepage</a>, with a dedicated page per sport such as <a href="https://orbistats.com/sports/tennis.html">tennis</a>. That is why the Sport type above has 13 members. Coverage can differ by product page, so verify each sport against the <a href="https://orbistats.com/developers/documentation.html">documentation</a> before you promise it to a tenant.</p>
<p>Now an implementation. Field names and paths are placeholders; compare a real payload from the public <a href="https://orbistats.com/developers/sandbox.html">API sandbox</a> and change only this file.</p>
<pre><code class="language-ts">// orbistats-provider.ts
import WebSocket from 'ws';
import { MatchEvent, SportsProvider, Sport } from './provider';

const REST_BASE = process.env.UPSTREAM_REST_BASE ?? 'https://api.orbistats.com/v1';
const WS_URL = process.env.UPSTREAM_WS_URL!;
const KEY = process.env.UPSTREAM_API_KEY!;

export class OrbistatsProvider implements SportsProvider {
  private handlers: Array&lt;(e: MatchEvent) =&gt; void&gt; = [];
  private attempt = 0;

  onEvent(h: (e: MatchEvent) =&gt; void) { this.handlers.push(h); }

  async snapshot(sport: Sport, resource: string, params: Record&lt;string, string&gt;) {
    const url = new URL(`${REST_BASE}/${sport}/${resource}`);
    for (const [k, v] of Object.entries(params)) url.searchParams.set(k, v);
    const res = await fetch(url, {
      headers: { Authorization: `Bearer ${KEY}` },
      signal: AbortSignal.timeout(8000),
    });
    if (!res.ok) throw new Error(`upstream ${res.status}`);
    return res.json();
  }

  start() {
    const ws = new WebSocket(WS_URL, { headers: { Authorization: `Bearer ${KEY}` } });
    let alive = true;
    let hb: NodeJS.Timeout;

    ws.on('open', () =&gt; {
      this.attempt = 0;
      // PLACEHOLDER: copy the real subscribe payload from the WebSocket docs
      ws.send(JSON.stringify({ action: 'subscribe', channel: 'live' }));
      hb = setInterval(() =&gt; {
        if (!alive) return ws.terminate();
        alive = false;
        ws.ping();
      }, 30000);
    });

    ws.on('pong', () =&gt; { alive = true; });

    ws.on('message', buf =&gt; {
      let msg: any;
      try { msg = JSON.parse(buf.toString()); } catch { return; }
      const items = Array.isArray(msg.data) ? msg.data : [msg.data ?? msg];
      for (const raw of items) {
        const e = this.toEvent(raw);
        if (e) this.handlers.forEach(h =&gt; h(e));
      }
    });

    ws.on('error', err =&gt; console.error('upstream error', err.message));
    ws.on('close', () =&gt; {
      clearInterval(hb);
      const delay = Math.min(30000, 1000 &lt;&lt; this.attempt++) + Math.floor(Math.random() * 500);
      setTimeout(() =&gt; this.start(), delay);
    });
  }

  private toEvent(raw: any): MatchEvent | null {
    if (!raw || typeof raw !== 'object') return null;
    const matchId = raw.fixture_id ?? raw.match_id ?? raw.id;
    if (!matchId || !raw.sport) return null;
    return {
      eventId: String(raw.event_id ?? `${matchId}:${Date.now()}`),
      matchId: String(matchId),
      sport: raw.sport,
      kind: raw.type ?? 'score',
      payload: raw,
      receivedAt: Date.now(),
    };
  }
}
</code></pre>
<p>Notice what this achieves. One upstream WebSocket serves every tenant. If a thousand tenants connect to you, the provider still sees one connection. That protects your plan limits and keeps your bill predictable. Read the <a href="https://orbistats.com/developers/changelog.html">changelog</a> before you ship, because streaming behaviour and API details can change.</p>
<p>Layer 2: Tenants, keys and plans</p>
<p>Model the tenant first, because every other component asks one question: who is calling, and what are they allowed to do?</p>
<pre><code class="language-sql">-- schema.sql
CREATE TABLE tenants (
  id uuid PRIMARY KEY,
  name text NOT NULL,
  plan text NOT NULL DEFAULT 'free',
  created_at timestamptz NOT NULL DEFAULT now()
);

CREATE TABLE api_keys (
  id uuid PRIMARY KEY,
  tenant_id uuid NOT NULL REFERENCES tenants(id),
  key_hash text NOT NULL UNIQUE,
  label text,
  created_at timestamptz NOT NULL DEFAULT now(),
  revoked_at timestamptz
);

CREATE TABLE usage_daily (
  tenant_id uuid NOT NULL REFERENCES tenants(id),
  day date NOT NULL,
  requests bigint NOT NULL DEFAULT 0,
  PRIMARY KEY (tenant_id, day)
);

CREATE TABLE tenant_webhooks (
  id uuid PRIMARY KEY,
  tenant_id uuid NOT NULL REFERENCES tenants(id),
  url text NOT NULL,
  secret text NOT NULL,
  sports text[] NOT NULL DEFAULT '{}'
);
</code></pre>
<p>Never store raw API keys. Generate a long random key, show it to the tenant once, and store only a hash. Because the key has high entropy, a fast hash like SHA-256 is fine.</p>
<pre><code class="language-ts">// keys.ts
import { createHash, randomBytes } from 'node:crypto';

export function generateKey(): { raw: string; hash: string } {
  const raw = 'sk_live_' + randomBytes(24).toString('hex');
  return { raw, hash: hashKey(raw) };
}

export function hashKey(raw: string): string {
  return createHash('sha256').update(raw).digest('hex');
}
</code></pre>
<p>Plans are configuration, not code branches. These are your own plans for your own tenants, and the numbers are illustrative:</p>
<pre><code class="language-ts">// plans.ts
import { Sport } from './provider';

export interface Plan {
  burst: number;
  msPerToken: number;
  dailyQuota: number;
  sports: Sport[] | 'all';
  realtime: boolean;
  delayMs: number;
  maxConnections: number;
}

export const PLANS: Record&lt;string, Plan&gt; = {
  free:       { burst: 10,  msPerToken: 500, dailyQuota: 1000,    sports: ['football', 'basketball', 'tennis'], realtime: false, delayMs: 45000, maxConnections: 1 },
  pro:        { burst: 50,  msPerToken: 100, dailyQuota: 100000,  sports: 'all', realtime: true,  delayMs: 0, maxConnections: 5 },
  enterprise: { burst: 200, msPerToken: 20,  dailyQuota: 5000000, sports: 'all', realtime: true,  delayMs: 0, maxConnections: 50 },
};

export function canAccess(plan: Plan, sport: Sport): boolean {
  return plan.sports === 'all' || plan.sports.includes(sport);
}
</code></pre>
<p>A tenant lookup happens on every request, so cache it briefly:</p>
<pre><code class="language-ts">// auth.ts
import { hashKey } from './keys';

export interface Tenant { id: string; plan: string; }

const cache = new Map&lt;string, { tenant: Tenant; expires: number }&gt;();

export async function authenticate(
  rawKey: string | undefined,
  db: { findTenantByKeyHash(h: string): Promise&lt;(Tenant &amp; { revokedAt: Date | null }) | null&gt; }
): Promise&lt;Tenant | null&gt; {
  if (!rawKey) return null;
  const hash = hashKey(rawKey);

  const hit = cache.get(hash);
  if (hit &amp;&amp; hit.expires &gt; Date.now()) return hit.tenant;

  const row = await db.findTenantByKeyHash(hash);
  if (!row || row.revokedAt) {
    cache.delete(hash);
    return null;
  }
  cache.set(hash, { tenant: { id: row.id, plan: row.plan }, expires: Date.now() + 30000 });
  return { id: row.id, plan: row.plan };
}
</code></pre>
<p>The trade-off is explicit: a revoked key can keep working for up to 30 seconds. If that is too long for your threat model, shorten the cache time or publish revocation events to all nodes.</p>
<p>Layer 3: Rate limiting that survives many servers</p>
<p>A per-process counter breaks the moment you run two servers. Use a shared token bucket in Redis, executed atomically in a Lua script so two nodes cannot race each other.</p>
<pre><code class="language-ts">// ratelimit.ts
import Redis from 'ioredis';

const redis = new Redis(process.env.REDIS_URL!);

const SCRIPT = `
local key = KEYS[1]
local capacity = tonumber(ARGV[1])
local msPerToken = tonumber(ARGV[2])
local now = tonumber(ARGV[3])
local cost = tonumber(ARGV[4])
local ttl = tonumber(ARGV[5])

local data = redis.call('HMGET', key, 'tokens', 'ts')
local tokens = tonumber(data[1])
local ts = tonumber(data[2])
if tokens == nil then
  tokens = capacity
  ts = now
end

local elapsed = math.max(0, now - ts)
tokens = math.min(capacity, tokens + elapsed / msPerToken)

local allowed = 0
if tokens &gt;= cost then
  tokens = tokens - cost
  allowed = 1
end

redis.call('HMSET', key, 'tokens', tokens, 'ts', now)
redis.call('PEXPIRE', key, ttl)
return { allowed, math.floor(tokens) }
`;

export async function takeToken(tenantId: string, burst: number, msPerToken: number, cost = 1) {
  const ttl = Math.ceil(burst * msPerToken) + 1000;
  const [allowed, remaining] = (await redis.eval(
    SCRIPT, 1, `rl:${tenantId}`, burst, msPerToken, Date.now(), cost, ttl
  )) as [number, number];
  return { allowed: allowed === 1, remaining };
}
</code></pre>
<p>The daily quota is a separate concern. Count in memory and flush in batches, so you do not write to the database on every request:</p>
<pre><code class="language-ts">// usage.ts
const pending = new Map&lt;string, number&gt;();

export function recordUsage(tenantId: string) {
  pending.set(tenantId, (pending.get(tenantId) ?? 0) + 1);
}

export async function flushUsage(db: { query(sql: string, params: unknown[]): Promise&lt;unknown&gt; }) {
  const batch = Array.from(pending.entries());
  pending.clear();
  for (const [tenantId, count] of batch) {
    await db.query(
      `INSERT INTO usage_daily (tenant_id, day, requests)
       VALUES ($1, CURRENT_DATE, $2)
       ON CONFLICT (tenant_id, day)
       DO UPDATE SET requests = usage_daily.requests + $2`,
      [tenantId, count]
    );
  }
}

setInterval(() =&gt; { /* call flushUsage(db) here */ }, 5000);
</code></pre>
<p>Two details here save you support tickets. Always return rate limit headers so tenants can self-correct, and return 429 with a Retry-After value instead of silently dropping requests.</p>
<p>Layer 4: The request pipeline</p>
<p>Order matters. Authenticate first, then check entitlements, then rate limit, then serve from cache. Cheap checks come before expensive ones.</p>
<pre><code class="language-ts">// api.ts
import express from 'express';
import { authenticate } from './auth';
import { PLANS, canAccess } from './plans';
import { takeToken } from './ratelimit';
import { recordUsage } from './usage';
import { swr } from './cache';
import { Sport } from './provider';

const app = express();

app.use('/v1', async (req, res, next) =&gt; {
  const raw = req.header('authorization')?.replace(/^Bearer /, '');
  const tenant = await authenticate(raw, db);
  if (!tenant) return res.status(401).json({ error: 'invalid_api_key' });

  const plan = PLANS[tenant.plan];
  const rl = await takeToken(tenant.id, plan.burst, plan.msPerToken);
  res.setHeader('X-RateLimit-Remaining', String(rl.remaining));
  if (!rl.allowed) {
    res.setHeader('Retry-After', String(Math.ceil(plan.msPerToken / 1000)));
    return res.status(429).json({ error: 'rate_limited' });
  }

  recordUsage(tenant.id);
  (req as any).tenant = tenant;
  (req as any).plan = plan;
  next();
});

app.get('/v1/:sport/:resource', async (req, res) =&gt; {
  const sport = req.params.sport as Sport;
  const plan = (req as any).plan;

  if (!canAccess(plan, sport)) {
    return res.status(403).json({ error: 'sport_not_in_plan' });
  }

  const key = `snap:${sport}:${req.params.resource}:${JSON.stringify(req.query)}`;
  try {
    const data = await swr(key, 15000, 120000, () =&gt;
      provider.snapshot(sport, req.params.resource, req.query as Record&lt;string, string&gt;)
    );
    res.json(data);
  } catch {
    res.status(502).json({ error: 'upstream_unavailable' });
  }
});
</code></pre>
<p>Validate the sport against your allow-list before building any upstream path. Otherwise your gateway becomes an open proxy.</p>
<p>Layer 5: Caching that protects the upstream</p>
<p>This is where a multi-tenant platform earns its margin. If 500 tenants ask for today's football fixtures, you should make one upstream request, not 500. Two techniques do the work: request coalescing and stale-while-revalidate.</p>
<pre><code class="language-ts">// cache.ts
const inflight = new Map&lt;string, Promise&lt;unknown&gt;&gt;();

export function coalesce&lt;T&gt;(key: string, loader: () =&gt; Promise&lt;T&gt;): Promise&lt;T&gt; {
  const existing = inflight.get(key);
  if (existing) return existing as Promise&lt;T&gt;;
  const p = loader().finally(() =&gt; inflight.delete(key));
  inflight.set(key, p);
  return p;
}

interface Entry { value: unknown; freshUntil: number; staleUntil: number; }
const store = new Map&lt;string, Entry&gt;();

export async function swr&lt;T&gt;(
  key: string,
  freshMs: number,
  staleMs: number,
  loader: () =&gt; Promise&lt;T&gt;
): Promise&lt;T&gt; {
  const now = Date.now();
  const hit = store.get(key);

  if (hit &amp;&amp; now &lt; hit.freshUntil) return hit.value as T;

  const refresh = () =&gt;
    coalesce(key, loader).then(value =&gt; {
      store.set(key, { value, freshUntil: Date.now() + freshMs, staleUntil: Date.now() + freshMs + staleMs });
      return value;
    });

  if (hit &amp;&amp; now &lt; hit.staleUntil) {
    refresh().catch(() =&gt; {});
    return hit.value as T;
  }
  return refresh();
}
</code></pre>
<p>Coalescing stops a cache stampede: when a hot key expires, only one caller goes upstream and everyone else waits for the same promise. Stale-while-revalidate means that if the upstream has a bad minute, tenants still get the last good answer instead of errors. Match the freshness window to the data. Standings can be cached far longer than live scores, and historical data far longer than both. For multiple nodes, move the same logic into Redis.</p>
<p>Layer 6: Real-time fan-out to tenants</p>
<p>The stream gateway takes events from the internal bus and sends each one only to tenants entitled to see it. Three problems need handling: entitlement filtering, delayed data for lower plans, and slow consumers.</p>
<pre><code class="language-ts">// gateway.ts
import { WebSocketServer, WebSocket } from 'ws';
import { authenticate } from './auth';
import { PLANS, canAccess } from './plans';
import { MatchEvent, Sport } from './provider';

interface Conn {
  ws: WebSocket;
  tenantId: string;
  sports: Set&lt;Sport&gt;;
  delayMs: number;
  dropped: number;
}

const MAX_BUFFERED = 1000000;
const MAX_DROPS = 200;
const conns = new Set&lt;Conn&gt;();
const perTenant = new Map&lt;string, number&gt;();

const wss = new WebSocketServer({ port: 4002, path: '/v1/stream' });

wss.on('connection', async (ws, req) =&gt; {
  const raw = req.headers['authorization']?.toString().replace(/^Bearer /, '');
  const tenant = await authenticate(raw, db);
  if (!tenant) return ws.close(4401, 'invalid_api_key');

  const plan = PLANS[tenant.plan];
  if (!plan.realtime &amp;&amp; plan.delayMs === 0) return ws.close(4403, 'realtime_not_in_plan');

  const open = perTenant.get(tenant.id) ?? 0;
  if (open &gt;= plan.maxConnections) return ws.close(4429, 'too_many_connections');
  perTenant.set(tenant.id, open + 1);

  const conn: Conn = { ws, tenantId: tenant.id, sports: new Set(), delayMs: plan.delayMs, dropped: 0 };
  conns.add(conn);

  ws.on('message', buf =&gt; {
    try {
      const msg = JSON.parse(buf.toString());
      if (msg.action === 'subscribe' &amp;&amp; Array.isArray(msg.sports)) {
        for (const s of msg.sports) {
          if (canAccess(plan, s)) conn.sports.add(s);
        }
        ws.send(JSON.stringify({ type: 'subscribed', sports: Array.from(conn.sports) }));
      }
    } catch { /* ignore malformed client frames */ }
  });

  ws.on('close', () =&gt; {
    conns.delete(conn);
    perTenant.set(tenant.id, Math.max(0, (perTenant.get(tenant.id) ?? 1) - 1));
  });
});

function deliver(conn: Conn, frame: string) {
  if (conn.ws.readyState !== WebSocket.OPEN) return;
  if (conn.ws.bufferedAmount &gt; MAX_BUFFERED) {
    conn.dropped++;
    if (conn.dropped &gt; MAX_DROPS) conn.ws.close(1013, 'slow_consumer');
    return;
  }
  conn.ws.send(frame);
}

export function publish(e: MatchEvent) {
  const frame = JSON.stringify({ type: e.kind, sport: e.sport, matchId: e.matchId, data: e.payload });
  for (const conn of conns) {
    if (!conn.sports.has(e.sport)) continue;
    if (conn.delayMs &gt; 0) setTimeout(() =&gt; deliver(conn, frame), conn.delayMs);
    else deliver(conn, frame);
  }
}
</code></pre>
<p>Wire it up in one place:</p>
<pre><code class="language-ts">provider.onEvent(publish);
provider.start();
</code></pre>
<p>The slow consumer guard deserves attention. If a tenant's client cannot keep up, its socket buffer grows and eats your memory. Checking bufferedAmount and eventually disconnecting that one client protects everyone else. This is the noisy neighbor problem in its purest form.</p>
<p>The delayed delivery in the free plan is a product decision you make for your own tenants. It is separate from whatever delay your upstream plan has, which is why you should read the provider's <a href="https://orbistats.com/pricing.html">pricing page</a> rather than assume that free-tier data is real time.</p>
<p>Layer 7: Tenant isolation in the database</p>
<p>Application code that filters by tenant_id everywhere works until someone forgets one WHERE clause. Row-level security makes the database enforce it:</p>
<pre><code class="language-sql">ALTER TABLE tenant_webhooks ENABLE ROW LEVEL SECURITY;

CREATE POLICY tenant_isolation ON tenant_webhooks
  USING (tenant_id = current_setting('app.tenant_id')::uuid);
</code></pre>
<p>In the application, set the tenant inside the same transaction as the query:</p>
<pre><code class="language-ts">// tenant-db.ts
export async function withTenant&lt;T&gt;(
  pool: { connect(): Promise&lt;any&gt; },
  tenantId: string,
  fn: (client: any) =&gt; Promise&lt;T&gt;
): Promise&lt;T&gt; {
  const client = await pool.connect();
  try {
    await client.query('BEGIN');
    await client.query(`SELECT set_config('app.tenant_id', $1, true)`, [tenantId]);
    const result = await fn(client);
    await client.query('COMMIT');
    return result;
  } catch (err) {
    await client.query('ROLLBACK');
    throw err;
  } finally {
    client.release();
  }
}
</code></pre>
<p>The third argument true makes the setting local to the transaction. That matters with connection pools, because a setting that leaks to the next request would be a data leak between tenants. Also connect as a database role that is not a superuser and does not own the tables, since owners can bypass row-level security by default.</p>
<p>Layer 8: Outbound webhooks done properly</p>
<p>Tenants will want push notifications too. Inbound, Orbistats documents its own <a href="https://orbistats.com/api/webhooks-api.html">webhooks</a>; outbound, you become the sender, and you inherit every reliability problem that implies. Sign every payload, include an id for idempotency, and retry with backoff.</p>
<pre><code class="language-ts">// webhooks-out.ts
import { createHmac, randomUUID } from 'node:crypto';

const SCHEDULE_MS = [0, 5000, 30000, 300000, 1800000];

export function sign(secret: string, timestamp: number, body: string): string {
  return createHmac('sha256', secret).update(`${timestamp}.${body}`).digest('hex');
}

export async function sendWebhook(url: string, secret: string, event: object) {
  const body = JSON.stringify({ id: randomUUID(), ...event });

  for (let attempt = 0; attempt &lt; SCHEDULE_MS.length; attempt++) {
    if (SCHEDULE_MS[attempt] &gt; 0) await new Promise(r =&gt; setTimeout(r, SCHEDULE_MS[attempt]));

    const ts = Math.floor(Date.now() / 1000);
    try {
      const res = await fetch(url, {
        method: 'POST',
        headers: {
          'Content-Type': 'application/json',
          'X-Signature': sign(secret, ts, body),
          'X-Timestamp': String(ts),
        },
        body,
        signal: AbortSignal.timeout(5000),
      });
      if (res.ok) return true;
      if (res.status &gt;= 400 &amp;&amp; res.status &lt; 500 &amp;&amp; res.status !== 429) return false;
    } catch { /* network error, retry */ }
  }
  return false;
}
</code></pre>
<p>Tell your tenants to verify the signature and reject old timestamps, and to deduplicate by event id. Also never follow redirects blindly and never let tenants point webhooks at your internal network addresses, or you have built a server-side request forgery machine.</p>
<p>Observability per tenant</p>
<p>Metrics without a tenant label are nearly useless in a multi-tenant system. At minimum record, per tenant and per sport: request count, error rate, rate-limited count, stream connections, dropped frames, and delivery lag.</p>
<p>Delivery lag is worth measuring yourself. Providers publish marketing figures, and the sub-50ms latency claim on the Orbistats site is the vendor's own statement, not an independent benchmark. Timestamp every event when it arrives from upstream and again when you write it to a tenant socket, and export the difference as a histogram. Then you know your own number instead of quoting anyone else's.</p>
<pre><code class="language-ts">// lag.ts
export function observeLag(e: { receivedAt: number }, tenantId: string, record: (t: string, ms: number) =&gt; void) {
  record(tenantId, Date.now() - e.receivedAt);
}
</code></pre>
<p>Failure modes and what to do about them</p>
<table>
<thead>
<tr>
<th>Failure</th>
<th>What tenants see</th>
<th>Design response</th>
</tr>
</thead>
<tbody><tr>
<td>Upstream WebSocket drops</td>
<td>Stale live data</td>
<td>Reconnect with backoff and jitter, mark feed status, serve cached snapshots</td>
</tr>
<tr>
<td>Upstream rate limit or 5xx</td>
<td>Errors on REST</td>
<td>Retry with backoff, stale-while-revalidate, coalesce requests</td>
</tr>
<tr>
<td>Cache stampede</td>
<td>Latency spikes</td>
<td>Single-flight coalescing per key</td>
</tr>
<tr>
<td>Noisy tenant</td>
<td>Everyone slows down</td>
<td>Per-tenant token bucket and connection caps</td>
</tr>
<tr>
<td>Slow stream consumer</td>
<td>Memory growth</td>
<td>Buffer checks, drop frames, disconnect</td>
</tr>
<tr>
<td>Malformed upstream payload</td>
<td>Crashes or garbage</td>
<td>Validate in the adapter, return null, never trust input</td>
</tr>
<tr>
<td>Leaked API key</td>
<td>Unauthorized use</td>
<td>Hashed storage, revocation, short auth cache, per-key labels</td>
</tr>
<tr>
<td>Tenant webhook endpoint down</td>
<td>Missing notifications</td>
<td>Signed retries, backoff schedule, dead-letter list</td>
</tr>
<tr>
<td>Provider changes an API</td>
<td>Sudden breakage</td>
<td>Adapter isolation, changelog monitoring, contract tests against the sandbox</td>
</tr>
</tbody></table>
<p>The licensing question nobody wants to talk about</p>
<p>This is the least fun section and possibly the most important. Sports data is licensed content. If you take an upstream feed and redistribute it to your own customers, your agreement with the provider has to allow that. Many standard plans allow you to use data inside your own application but not to resell or re-serve raw feeds to third parties.</p>
<p>Before building anything in this post for a real product:</p>
<ul>
<li>Read the provider's terms on redistribution and sublicensing</li>
<li>Ask about enterprise or custom arrangements. Orbistats positions enterprise plans around custom feeds, dedicated infrastructure and SLA requirements, which is the kind of conversation you want to have early</li>
<li>Keep attribution and usage restrictions in your own tenant terms</li>
<li>Do not assume that technical access equals legal permission</li>
</ul>
<p>Design lessons</p>
<p>These are the patterns that shaped the design above, presented as principles rather than war stories:</p>
<ol>
<li>Put a wall between the provider and your product. One adapter file, one internal event shape, and no vendor JSON leaking into business logic.</li>
<li>Share the expensive thing, isolate the cheap thing. One upstream subscription, but per-tenant limits, entitlements and data.</li>
<li>Cheap checks first. Authenticate, then entitlements, then rate limit, then cache, and only then touch the upstream.</li>
<li>Cache with intent. Coalescing and stale-while-revalidate protect both your upstream quota and your tenants from upstream bad days.</li>
<li>Assume slow clients exist. Every fan-out system needs a slow consumer policy before the first slow consumer shows up.</li>
<li>Make the database enforce isolation. Row-level security catches the query someone forgot to filter.</li>
<li>Label everything by tenant. You cannot fix a noisy neighbor you cannot see.</li>
<li>Measure your own latency. Never publish a speed claim you have not measured on your own path.</li>
<li>Treat licensing as architecture. It decides what you are allowed to build.</li>
<li>Test against the sandbox. Contract tests using real sample responses catch provider changes early, and the provider's <a href="https://orbistats.com/developers/changelog.html">changelog</a> tells you what to look for.</li>
</ol>
<p>What to build next</p>
<ul>
<li>Add a historical backfill service using the <a href="https://orbistats.com/api/historical-sports-data-api.html">Historical Sports Data API</a> so new tenants get archives without touching live infrastructure.</li>
<li>Add an odds product tier using the <a href="https://orbistats.com/api/odds-api.html">Odds API</a>, with its own entitlements, since market depth can vary by competition.</li>
<li>Add per-sport semantics. Cricket, tennis, golf and horse racing do not map cleanly to home and away, so extend the event model per sport.</li>
<li>Add multi-region failover once one region is stable.</li>
<li>Document your own API well. The <a href="https://orbistats.com/developers.html">developer hub</a> and <a href="https://orbistats.com/developers/documentation.html">documentation</a> at Orbistats are a good reference for what developers expect: quickstart, examples, sandbox and a changelog.</li>
<li>Read the <a href="https://orbistats.com/company/about.html">About page</a> to understand the provider's positioning before you build a business on top of any upstream.</li>
</ul>
<p>Wrap-up</p>
<p>A multi-tenant sports data platform is mostly not about sports. It is about identity, limits, isolation, caching and failure handling, wrapped around a feed you do not control. Build the wall around the provider first, share the upstream connection, enforce every limit at the edge, and let the database back up your isolation rules.</p>
<p>If you are just starting, do not build all of this at once. Explore the real response shapes in the <a href="https://orbistats.com/developers/sandbox.html">sandbox</a>, check what is available on the <a href="https://orbistats.com/pricing.html">pricing page</a>, and start with the adapter, the token bucket and the cache. Everything else can follow once real tenants show you where it hurts.</p>
<p>If you have run something like this in production, I would love to hear which layer surprised you the most.</p>
]]></content:encoded></item></channel></rss>