# Data sources

[Dataset](/docs/data) is the file of hidden values a round is scored
against. This chapter covers where those values come from: how a source is
turned into that file, who makes the request, and what happens when a step
fails.

## The pipeline

```
  A SOURCE            a feed you call, a feed we call, or a file you write
     |
     v
  TRANSFORM           one function per source. It maps the source's fields
                      onto our measure keys and normalises the units.
     |
     v
  A DATASET           the same shape, whatever the source was
     |
     v
  PUBLISHED           you push it, into a private prefix no browser reads
     |
     v
  A PRODUCER'S ROUND  the words a fan sees, each choice carrying a key
     |
     v
  JOINED              when the round is built, on that key
     |
     v
  THE ENGINE          holds the values. The browser gets labels only.
```

Every source ends in the same shape. Everything after the transform is the same
code, whatever the values came from.

Once a source is connected and pulling, [`gamestage round seed`](/docs/cli#round-seed-game-count-n)
builds candidate rounds straight from it and creates them in Studio, rather
than you typing a board by hand.

## Three parts

**A transform** is one function per source. It takes a payload and returns a
dataset. It does no fetching. `supplier-fpl.ts` and `supplier-nba.ts` in
`packages/schema` are the two we have.

**A measure key** is our name for a quantity, such as `player_price` or
`career_points`. The transform maps the source's field onto it and normalises
the units. Fantasy Premier League prices stay in tenths of a million, because a
dataset holds whole numbers and dividing here would round a budget a fan
can lose by a penny.

**An entry key** is what a producer types against each choice in Studio, in the
`entry_key` field. It joins a label a fan reads to a value they never see. Use
the source's own id, not a name. Names repeat, differ between feeds, and change.

## Measure keys are not a shared vocabulary

There is no canonical list of measure keys. Each transform holds its own mapping
table and each source declares what it produces, but nothing checks a key
against a list. Two sources can use one key for different quantities.

This is safe while a workspace's dataset comes from one source, because a
round names the measure it is counted on. It stops being safe when one game
draws on two sources. Tell us before you build that, not after.

## Four sources

The register is an array in `packages/schema/src/suppliers.ts`. Adding a source
is a code change and a release.

| Source | Standing | Who calls | Needs a key | Available to |
| --- | --- | --- | --- | --- |
| Your own data | creator | you | No | Everyone, on every plan |
| Fantasy Premier League | unofficial | us | No | Everyone, on every plan |
| ESPN | unofficial | us | No | Everyone who switches it on |
| Sportmonks football | creator | us, with your key | Yes | Everyone who connects a Sportmonks key |

The table says who may use a source, not what runs today. `gamestage sources`
prints the same list with its state for your workspace: connected, available,
available once you add a key, or not available with the reason.

Your own data works now. Fantasy Premier League runs now: a scheduled pull
fetches every Premier League player's price, season points and points per
game once a day, with each player's name, club and position, so a producer
picks players by name rather than by id. The Points sample game plays on it.
ESPN runs once you [switch it on](#switch-on-espn): a scheduled pull fetches
season statistics for every player in the league you choose, twice a day.
Sportmonks runs once you [connect your own key](#connect-your-own-key): a
scheduled pull fetches season statistics for whichever competitions your plan
covers, with each player's name, team and position.

What a creator sees before connecting:

- **Your own data**: a file you push with `gamestage reference push` from the
  machine you signed in on.
- **Fantasy Premier League**: free, with no contract behind it. Prices move
  weekly, and a round in play keeps the ones it opened with.
- **ESPN**: free, with no contract behind it. Player season statistics for
  the NBA, WNBA, NFL, MLB and NHL. It carries no Premier League or other
  association football.
- **Sportmonks football**: your own Sportmonks plan and key. Their lowest,
  no-cost plan covers a few smaller leagues; the Premier League needs a paid
  one. Check your plan's terms cover this use before a game with a prize.

A source you cannot use is still listed, with the reason. Hiding it would mean
you never learn it exists.

## Switch on ESPN

ESPN needs no key. Switch it on for your workspace, one league at a time:

```
gamestage sources connect espn --league nba --measures points,rebounds,assists
```

Or open **Datasets** in Stage, find ESPN under Sources and choose **Switch
on**: pick a league, tick the measures you want, and name the dataset.

| Option | Type | Required | Default | What it does |
| --- | --- | --- | --- | --- |
| `--league` | `nba`, `wnba`, `nfl`, `mlb` or `nhl` | Yes | none | The league to fetch |
| `--measures` | comma-separated measure keys | No | a starting set for the league | The numbers a round can be scored on |
| `--season` | year | No | the current season | Which season's statistics |
| `--name` | dataset name | No | `espn-<league>` | What a game names in Studio to play on it |

`gamestage sources status espn --league nba` lists the measure keys a league
offers. `gamestage sources status espn` lists what you have switched on and
how often each is fetched.

The first fetch comes within the hour. After that it is fetched every 12
hours, and you cannot ask for more often: if a game needs fresher numbers, ask
Monterosa, who can set a shorter interval for one dataset.

`gamestage sources disconnect espn --name espn-nba` stops the fetching. The
data already fetched stays, so a game playing on it keeps working.

ESPN is unofficial. Read [what that means for you](#what-an-unofficial-source-means-for-you)
before building on it: use it to build and test, and do not launch a public
game on it without the data owner's permission.

A switch-on the CLI refuses says why: a league ESPN does not cover, a measure
the league does not offer, or a dataset name already used by another source.

## Connect your own key

Sportmonks needs your own key, because the arrangement is yours: your plan,
your terms, your limits. Connect it without it passing through a chat or a
terminal:

```
gamestage sources connect sportmonks
```

That prints a link to the Data area in Stage, where you paste the key into a
masked field. It is checked with Sportmonks and stored in AWS Secrets Manager,
never shown again. `gamestage sources status sportmonks` says whether it is
connected and what your plan covers, and `gamestage sources disconnect
sportmonks` removes it; the datasets you already have stay.

Once it is connected, we call Sportmonks on a schedule using your key: your
plan's limits are the ones spent, and your terms are what bind the call. What
is ours is the machine the schedule runs on. [The CLI reference](/docs/cli) has
the full command options, including reading the key from your own 1Password at
your own terminal with `--from-op`, which an agent must never run for you.

## Standing

Standing describes the arrangement behind a source, not the quality of its
data.

- **`licensed`**: a contract exists. Suitable for a game with something at stake.
- **`unofficial`**: no contract, no support, and it can change without notice.
  Suitable for a prototype.
- **`creator`**: your own data, or your own licensed feed. Your rights, your
  call.

Fantasy Premier League looks licensed and is not. Nothing is promised about that
endpoint, and a fantasy price is a number from their game rather than a
competition record. It stays `unofficial`.

### What an unofficial source means for you

Every unofficial source, Fantasy Premier League today, carries this, and
`gamestage sources`, `verify`, `deploy` and `promote` all say it:

> Requests to this source are made on behalf of you, the game builder. You are responsible for the data you consume, and for securing the legal rights and licence to use it. We recommend using it for testing only, and not launching with it unless you have explicit permission from the data owner.

It is a recommendation, so nothing refuses a game over it. It is said at each
of those steps because each is a step closer to fans.

## Bring your own licensed feed

The `creator` standing covers a feed you license as well as a file you write.
The arrangement is yours: your contract, your terms, your key. Gamestage is not
a party to it, which is the property a rights-holder with a data licence needs.
The values are fetched server-side and held by the Engine, and no response a
browser receives carries them, so the game never exposes what the licence
protects.

What that means today: you transform the feed into a dataset and push it with
`gamestage reference push`, on whatever schedule your source updates.

Open and public data sources sit under `unofficial`, each with its own licence
to respect. They are right for a prototype or a taster. A game with something
at stake needs a source with a contract behind it.

## Who makes the request

Each source also records who makes the call.

A source we call is bound by terms that bind us, and if they refuse us it
affects every game on the platform. A source you call is bound by your terms,
runs from your machine, and cannot affect anybody else.

Sportmonks is a third shape: we make the call, on a schedule, but with your
key. Your plan's terms and limits are the ones spent; what is ours is the
machine the schedule runs on. Disconnecting your key stops the pull without
affecting anybody else's game.

This is why changing source later is a row in the register rather than a
rewrite, and why pushing your own file is the lowest-risk option.

## Choosing a source

Decide this before you design the game.

**Your own data.** A club knows its squad, a broadcaster knows its schedule, a
publisher owns its archive. No licence and no expiry.

**A feed you already pay for.** Most broadcasters hold a match-data licence, so
the work is a transform of data you are entitled to. Check that your licence
covers this use.

**A provider you would need to license.** A real cost and a real lead time.
Find out now, not after the game is built.

**Anything scraped.** We cannot help, and you should assume a game built on it
cannot carry a prize.

### Player wages, as an example

Wages look like an easy game and are a hard case. They are usually not in the
match-data feeds a broadcaster licenses, and the published figures are estimates
compiled by third parties under their own terms. Two consequences:

- A prize decided by an estimate can be disputed. If a fan can argue with the
  number, they can argue with the result.
- A licence that forbids redistribution may forbid this use, even though the
  values never reach a fan.

Design for the data you have. "Whose career total is higher" needs no feed.
"Whose wage is higher" needs a licence and a refresh.

## Errors

Each failure below has a defined behaviour.

- **A dataset names a source that does not exist.** Refused, and the source is
  named. A run that silently did four of five looks like a run that had four to
  do.
- **A workspace may not use the source.** Refused before the request is made, so
  it costs the source no traffic. The same check decides what Stage lists and
  what a write allows, so the two cannot disagree.
- **A fetch fails.** No data is written; the pointer is marked as failing.
  What is published stays published, and a game being played carries on with
  the values it has.
- **An entry has no value.** It is named and left out, never set to zero. A
  player worth nothing is a free pick in a game about a budget.
- **The data has not changed.** The digest-keyed copy is not rewritten; the
  pointer's timestamp moves so the next run knows this dataset was checked,
  rather than looking permanently overdue.
- **What we fetched is not a valid dataset.** It is validated before it is
  compared, so a run cannot report "unchanged" about bytes it never checked.
- **Two measures share a key.** Refused. A round names the measure it is counted
  on, so two measures under one key cannot say what scored it.
- **A choice has a key the file does not hold.** The whole round is refused and
  the missing keys are named. A choice silently worth zero looks like a hard
  round rather than a broken one.
- **Dataset changes while a round is live.** The round is left alone and
  the reason is logged against the game once. Fans keep the values they
  started on, and a deploy or a refresh applies the change.
- **The digest-keyed copy is immutable by convention.** Nothing in the publish
  path rewrites one, and the copy is written before the pointer, so a name cannot
  resolve to bytes nobody can fetch. Whether the bucket would refuse an overwrite
  is a storage setting rather than a promise made here.

## Adding a source we do not have

There are two options, and neither is a screen.

Write the file yourself and push it. This is the `creator` row. It needs nothing
from us and is on every plan.

Or we write a transform: a row in the register and a function that returns a
dataset. That is a code change and a release. There is no self-serve path
onto a new feed. Fantasy Premier League runs for every workspace because the
relationship with the Premier League is ours; a new source on that footing is
the same, a code change and a release. A source that runs on your own key,
such as Sportmonks, is different: you connect it yourself with `gamestage
sources connect`, and the scheduled pull starts once it is connected. Ask us
before building on a source neither route covers.

## Managing what you have

Pushing a dataset is CLI-only, with `gamestage reference push`. The Data
screen in Stage lists the sources available to your workspace and, in its own
table, the datasets a workspace already holds: name, source, which games use
it, its measures, and when it was last checked. Opening a row shows more
detail. It never shows a value.

Signing in is a device grant confirmed in a browser, and the session stays on
that machine, so a push from CI runs with a session placed there rather than
by signing in. Run the push on your own schedule: you know when your source
updates and we do not, the credentials are yours, and one command in your CI
is smaller than a scheduler somebody has to operate for you.
