Point-in-Time Language Models
Language models that know only what was knowable at a given date — and the research that builds, tests, and uses them.
A conventional language model is trained on a temporally indiscriminate pile of text. Ask it to “forecast” 2015 from 2014's news and it quietly cheats: it has already read 2016. A point-in-time (or chronologically consistent) language model is trained only on text available up to a fixed cut-off date. A sequence of such models, one per year, quarter, or month, lets you replay history without look-ahead bias.
The idea began in finance, where leakage invalidates a backtest, but the same construction is a time capsule for any field in which knowledge changes: medicine, law, journalism, policy, and the social sciences. This site collects the model sequences that exist, the corpora one could train them on, the literature on look-ahead bias and temporal generalisation, and plans for what to build next.
Cut the dated stream at t₁ < t₂ < t₃, train one model per cut, and evaluate each model only on the interval after its cut-off. Leakage is removed structurally, not by prompting.
Model sequences
Six open efforts have released sequences of date-stamped checkpoints. They differ in domain, cadence, and in whether the first checkpoint was itself trained from scratch on period text or adapted from a modern base model (which reintroduces some leakage). Details, checkpoint lists, and a coverage chart are on the models page.
| Family | Architecture | Coverage | Cadence | Corpus |
|---|---|---|---|---|
| ChronoBERT / ChronoGPT | BERT-style encoder; GPT-style decoder; instruct variant | 1999–2024 | yearly (26 + 26 + 26) | chronologically filtered web and news text |
| Scaling PiT LMs | decoder-only, up to 4B params | 2013–2024 | monthly | 1T chronologically filtered FineWeb tokens |
| Time Machine GPT | GPT-2 | 2011–2022 | yearly (12) | WMT News Crawl + Wikipedia, datasets released |
| TimeLMs | RoBERTa-base (and one large) | 2019-Q4–2022-Q4 | quarterly | |
| StoriesLM | BERT-style, expanding windows | 1900–1963 | yearly (64) | American Stories newspapers |
| HistBERT | BERT-base, continued pretraining | 1900s–2000s | decadal (10) | COHA |
Research threads
The literature clusters into a few threads. The literature map draws the connections; the timeline dates them.
Building the models
Train a sequence of checkpoints on chronologically filtered text. Recent work shows the performance gap to unconstrained models narrows with scale.
Measuring look-ahead bias
How much of an LLM's “forecasting skill” is memorised outcome? Detection statistics, fake-date tests, and finance-specific benchmarks.
Can prompting substitute?
Telling a model to “pretend it is 2015” is cheap. It works for direct queries and fails for causally downstream knowledge; models struggle with chronology itself.
Temporal generalisation
The NLP thread that preceded finance: models degrade as the world moves past their training window, and continual pretraining only partly helps.
Temporal knowledge benchmarks
What does a model know as of when? Cutoff tracing, chronological knowledge across domains, evolving medical guidelines.
Beyond finance
Historical LLMs as simulated informants for behavioural science, and proposed uses in medicine, law, history, journalism, and policy.
Corpora
Every checkpoint sequence is only as good as its dated text. The corpora page catalogues candidate sources by domain, with coverage, licensing, and the finest temporal granularity each supports, from Chronicling America and COHA to SEC EDGAR, Hansard, and Common Crawl derivatives such as FineWeb.
What to build next
The plans page holds three costed proposals for extending the open model and corpus sequences on a small budget: a focused $500k programme, an open-source corpus programme, and a decentralised Bittensor-subnet variant. They are working documents, not commitments.
Further reading
- He, Lv, Manela & Wu (2025). Chronologically Consistent Large Language Models — the paper that made the construction mainstream, with ChronoBERT and ChronoGPT.
- Kelly, Malamud, Schwab & Xu (2026). Scaling Point-in-Time Language Models — the gap to unconstrained models closes with scale.
- Ludwig, Mullainathan & Rambachan (2024). Large Language Models: An Applied Econometric Framework — why training leakage matters for inference, not just prediction.
- Point-in-Time LLMs Beyond Finance — the essay on this site: what temporally faithful models make possible in six other fields.
- Overview slides — a short presentation of the idea.
Bibliography
Grouped by thread. Every entry also appears on the literature map and the timeline. arXiv identifiers were checked against the arXiv API; SSRN and journal items against Crossref.
Model sequences
- He, S., Lv, L., Manela, A., and Wu, J. (2025). Chronologically Consistent Large Language Models. arXiv:2502.21206. ChronoBERT and ChronoGPT: 26 yearly checkpoints each, 1999–2024, trained only on text available at each date; competitive with BERT-class models on NLP benchmarks.
- He, S., Lv, L., Manela, A., and Wu, J. (2025). Instruction Tuning Chronologically Consistent Language Models. arXiv:2510.11677. A chat-style, instruction-tuned ChronoGPT per cut-off year, giving a conservative lower bound on forecast accuracy once leakage is removed.
- Kelly, B., Malamud, S., Schwab, J., and Xu, T. A. (2026). Scaling Point-in-Time Language Models. arXiv:2607.11889. Decoder-only models up to 4B parameters on 1T chronologically filtered FineWeb tokens; monthly checkpoints 2013–2024; the gap to open-weight frontier models narrows with scale.
- Drinkall, F., Rahimikia, E., Pierrehumbert, J. B., and Zohren, S. (2024). Time Machine GPT. Findings of NAACL 2024. arXiv:2404.18543. Twelve yearly “nonprognosticative” GPT-2 models (2011–2022) trained from scratch on WMT News Crawl and Wikipedia; models and training data on Hugging Face.
- Loureiro, D., Barbieri, F., Neves, L., Espinosa Anke, L., and Camacho-Collados, J. (2022). TimeLMs: Diachronic Language Models from Twitter. ACL 2022 demo. Quarterly RoBERTa checkpoints from 2019-Q4 to 2022-Q4, with a library that routes each tweet to the model that could have seen it.
- Sarkar, S. K. (2024). StoriesLM: A Family of Language Models With Time-Indexed Training Data. SSRN 4881024. Sixty-four yearly models, 1900–1963, each continuing the previous year's checkpoint on that year's American Stories newspaper text.
- Qiu, W., and Xu, Y. (2022). HistBERT: A Pre-trained Language Model for Diachronic Lexical Semantic Analysis. arXiv:2202.03612. Decade-wise BERT models continued on COHA, for tracking lexical semantic change.
- Röttger, P., and Pierrehumbert, J. B. (2021). Temporal Adaptation of BERT and Performance on Downstream Document Classification. Findings of EMNLP 2021. Early evidence on what temporal adaptation does and does not buy.
Look-ahead bias and training leakage
- Lopez-Lira, A., and Tang, Y. (2023). Can ChatGPT Forecast Stock Price Movements? arXiv:2304.07619. The result that prompted the leakage question: strong in-sample “forecasts” from a model that had read the outcomes.
- Glasserman, P., and Lin, C. (2023). Assessing Look-Ahead Bias in Stock Return Predictions Generated By GPT Sentiment Analysis. arXiv:2309.17322. Anonymising firms and dates to separate genuine sentiment signal from memorised outcomes.
- Sarkar, S. K., and Vafa, K. (2024). Lookahead Bias in Pretrained Language Models. SSRN 4754678. Direct measurement of leakage in pretrained models used for economic prediction.
- Rahimikia, E., and Drinkall, F. (2024). Re(Visiting) Large Language Models in Finance. SSRN 4963618. Financial applications revisited with point-in-time models in hand.
- Ludwig, J., Mullainathan, S., and Rambachan, A. (2024). Large Language Models: An Applied Econometric Framework. arXiv:2412.07031. When LLM outputs can and cannot be used as data in empirical research; training leakage is one of the two central problems.
- Gao, Z., Jiang, W., and Yan, Y. (2025). Detecting Lookahead Bias in LLM Forecasts. arXiv:2512.23847. A “lookahead propensity” statistic from date-only recall queries; it collapses to zero after the training cutoff.
- Eliseev, A., and Seleznev, S. (2026). Fake Date Tests: Can We Trust In-sample Accuracy of LLMs in Macroeconomic Forecasting? arXiv:2601.07992. Prompt-sensitivity tests; none of the tested models passed.
- Benhenda, M. (2026). Look-Ahead-Bench: a Standardized Benchmark of Look-ahead Bias in Point-in-Time LLMs for Finance. arXiv:2601.13770. Alpha decay across market regimes as the measure; standard LLMs show it, point-in-time models do not.
- Kong, Y., Lee, H., Hwang, Y., Lopez-Lira, A., Levy, B., Mehta, D., Wen, Q., Choi, C., Lee, Y., and Zohren, S. (2026). Evaluating LLMs in Finance Requires Explicit Bias Consideration. arXiv:2602.14233. Position paper: five biases including look-ahead; a review of 164 papers finds none discussed in more than 28% of them.
- Li, W. W., Wang, M., and Ma, T. (2026). Summoning the Oracle to Slay It: Mitigating Look-Ahead Bias in Financial Backtesting with LLMs. arXiv:2605.24564. “Parametric look-ahead bias” and an inference-time decoding fix.
- Cai, X., et al. (2026). OpenPM: Auditable Point-in-Time Evaluation for LLM Portfolio-Management Agents. arXiv:2608.09988. Every record visible to the agent must be available at decision time; runs emit a contamination certificate.
Prompted cut-offs and chronology
- Gao, X., Zhang, R., Du, D., Mahindre, S., Somayajula, S. A., and Xie, P. (2025). Can Prompts Rewind Time for LLMs? Evaluating the Effectiveness of Prompted Knowledge Cutoffs. arXiv:2510.02340. Prompted cut-offs work for directly queried facts and fail for semantic shift and causally related knowledge.
- Asai, M., et al. (2026). Can LLMs Be Constrained to the Past? Improving Knowledge Cutoff through Recall-Based Prompting. arXiv:2606.05804. Recall-based prompting narrows, but does not close, the gap.
- Wongchamcharoen, P. K., and Glasserman, P. (2025). Do Large Language Models (LLMs) Understand Chronology? arXiv:2511.14214. Frontier models preserve local order but fail to maintain a globally consistent timeline, undermining prompt-based defences.
- Cheng, J., Marone, M., Weller, O., Lawrie, D., Khashabi, D., and Van Durme, B. (2024). Dated Data: Tracing Knowledge Cutoffs in Large Language Models. arXiv:2403.12958. The effective cutoff differs from the reported one, and differs by resource.
- Zhao, B., Brumbaugh, Z., Wang, Y., Hajishirzi, H., and Smith, N. A. (2024). Set the Clock: Temporal Alignment of Pretrained Language Models. arXiv:2402.16797. Aligning a model's internal “now” to a target year.
Temporal generalisation and misalignment
- Lazaridou, A., et al. (2021). Mind the Gap: Assessing Temporal Generalization in Neural Language Models. NeurIPS 2021. Performance degrades on text from after the training period; the case for time-stratified evaluation.
- Luu, K., Khashabi, D., Gururangan, S., Mandyam, K., and Smith, N. A. (2021). Time Waits for No One! Analysis and Challenges of Temporal Misalignment. arXiv:2111.07408.
- Agarwal, O., and Nenkova, A. (2021). Temporal Effects on Pre-trained Models for Language Processing Tasks. arXiv:2111.12790.
- Dhingra, B., Cole, J. R., Eisenschlos, J. M., Gillick, D., Eisenstein, J., and Cohen, W. W. (2021). Time-Aware Language Models as Temporal Knowledge Bases. TACL. Prefix each training example with its date; the model learns to condition on time.
- Jang, J., et al. (2021). Towards Continual Knowledge Learning of Language Models. ICLR 2022.
- Jin, X., et al. (2021). Lifelong Pretraining: Continually Adapting Language Models to Emerging Corpora. arXiv:2110.08534.
- Zhang, M. J. Q., and Choi, E. (2023). Mitigating Temporal Misalignment by Discarding Outdated Facts. arXiv:2305.14824. Predict how long a fact stays true.
- Shin, C., et al. (2025). TARDIS: Mitigating Temporal Misalignment via Representation Steering. arXiv:2503.18693.
Temporal knowledge benchmarks
- Park, Y., Yoon, C., Park, J., Lee, D., Jeong, M., and Kang, J. (2024). ChroKnowledge: Unveiling Chronological Knowledge of Language Models in Multiple Domains. ICLR 2025. ChroKnowBench spans general, biomedical, legal, commonsense, and mathematical knowledge, separating evolving from constant facts. Code.
- Jang, J., et al. (2022). TemporalWiki: A Lifelong Benchmark for Training and Evaluating Ever-Evolving Language Models. EMNLP 2022. Code.
- Liška, A., et al. (2022). StreamingQA: A Benchmark for Adaptation to New Knowledge over Time in Question Answering Models. ICML 2022.
- Kasai, J., et al. (2022). RealTime QA: What's the Answer Right Now? arXiv:2207.13332.
- Deng, H., Jiao, W., Liu, X., Zhang, M., and Tu, Z. (2024). NewTerm: Benchmarking Real-Time New Terms for Large Language Models with Annual Updates. arXiv:2410.20814.
- Guan, Z., et al. (2026). Large Language Models Lack Temporal Awareness of Medical Knowledge. arXiv:2605.13045. TempoMed-Bench: evolving clinical guidelines as the test of “knowing when knowledge is correct”.
Corpora and data
- Dell, M., et al. (2023). American Stories: A Large-Scale Structured Text Dataset of Historical U.S. Newspapers. NeurIPS 2023 Datasets. Article-level text from nearly 20 million Chronicling America scans; the training data for StoriesLM.
- Yang, H., Liu, X.-Y., and Wang, C. D. (2023). FinGPT: Open-Source Financial Large Language Models. arXiv:2306.06031. Data-centric financial LLM pipeline; not point-in-time, but a source of dated financial text.
- See the corpora page for the full catalogue of dated text sources.
Beyond finance
- Varnum, M. E. W., Baumard, N., Atari, M., and Gray, K. (2024). Large Language Models based on historical text could offer informative tools for behavioral science. PNAS 121(42). Historical LLMs as simulated informants for the psychology of past populations.
- Point-in-Time LLMs Beyond Finance. This site. Medicine, law, history, journalism, policy, and the social sciences: what becomes testable when the model cannot see ahead.
Working on a point-in-time model, corpus, or benchmark that is missing here? Open an issue or pull request on microprediction/pitllm.