Plans

Three costed ways to extend the open point-in-time model and corpus sequences on a small budget.

The existing sequences were built by a handful of academic groups. The gaps are known: no yearly coverage of 1964–1998, no monthly cadence before 2013, no non-English or domain-specific sequence, and no shared evaluation of temporal consistency outside finance. The documents below are working proposals for closing some of these gaps with roughly half a million dollars or less. They differ mainly in who does the work and who pays for the data.

PlanBudgetDataWho does the workMain deliverable
Focused programme$500k, 12 monthsMixed: open sources plus licensed financial newsSmall salaried team, 2–3 peopleExtended TimeLMs and ChronoBERT sequences, a financial temporal benchmark, a backtesting tool
Open-source corpus programme~$550k, 12 months100% public domain or open licenceSmall team, 4–6 part-time rolesThree flagship dated corpora and open temporal benchmarks
Bittensor subnet$50k setup; work paid in emissions100% public domain or open licenceIncentivised miners, validated by the subnetThe same three corpora, produced by a decentralised network

A focused $500k programme

Full document. The plan opens with a reality check: $500k buys six to twelve months of a two- or three-person team, not two years of infrastructure. It therefore rules out training large models from scratch, building comprehensive data infrastructure, or producing production applications, and commits to extending what exists.

The financial-domain focus is a deliberate narrowing: it is where look-ahead bias has a measurable cost and where the evaluation literature already exists.

An open-source corpus programme

Full document. Rather than extending models, this plan builds the dated text they would need, using only public domain and openly licensed sources so that the outputs can be redistributed and maintained by the community.

Targets are stated as OKRs: 500GB–1TB of dated text across three domains, three to five years of added coverage per domain, and 98% temporal-consistency validation. The document's own team costing comes to more than the headline phase totals once overhead and contingency are added, which is worth knowing before quoting a single number.

A Bittensor subnet

Full document. The same three corpora and benchmarks as the open-source programme, but produced by miners on a Bittensor subnet and scored by validators, so the direct cost is the subnet setup rather than salaries. Miner tasks are temporal filtering, deduplication, quality scoring, and benchmark construction over the open sources; validators check temporal consistency and reward accordingly.

MetricCentralisedOpen-source subnet
Set-up cost~$742k$50k
Team6.5 FTE, one location50–100 miners, global
IncentiveSalariesToken emissions
Data licensingMixed100% open
After month 12Project endsNetwork continues

The document's emissions figures assume a fixed TAO price and a fixed share of network emissions; both are volatile, and the “$0 direct cost” of phases 2 and 3 is cost borne by the network rather than cost that disappears. Treat the comparison as directional.

What would change the picture

These are working documents. Suggestions and alternative costings are welcome on GitHub.