Point-in-Time Language Models
Language models that know only what was knowable at a given date — and the research that builds, tests, and uses them.
A conventional language model is trained on a temporally indiscriminate pile of text. Ask it to “forecast” 2015 from 2014's news and it quietly cheats: it has already read 2016. A point-in-time (or chronologically consistent) language model is trained only on text available up to a fixed cut-off date. A sequence of such models, one per year, quarter, or month, lets you replay history without look-ahead bias.
The idea began in finance, where leakage invalidates a backtest, but the same construction is a time capsule for any field in which knowledge changes: medicine, law, journalism, policy, and the social sciences. This site collects the model sequences that exist, the corpora one could train them on, the literature on look-ahead bias and temporal generalisation, and plans for what to build next.
Cut the dated stream at t₁ < t₂ < t₃, train one model per cut, and evaluate each model only on the interval after its cut-off. Leakage is removed structurally, not by prompting.
Model sequences
Seven open efforts have released sequences of date-stamped checkpoints, and TiMoE serves every cut-off from one set of routed experts. They differ in domain, cadence, and in whether the first checkpoint was itself trained from scratch on period text or adapted from a modern base model (which reintroduces some leakage). Details, checkpoint lists, and a coverage chart are on the models page.
| Family | Architecture | Coverage | Cadence | Corpus |
|---|---|---|---|---|
| ChronoBERT / ChronoGPT | BERT-style encoder; GPT-style decoder; instruct variant | 1999–2024 | yearly (26 + 26 + 26) | chronologically filtered web and news text |
| Scaling PiT LMs | decoder-only, up to 4B params | 2013–2024 | monthly | 1T chronologically filtered FineWeb tokens |
| DatedGPT | decoder-only, 1.3B params; instruct variant | 2013–2024 | yearly (12) | ~100B tokens per year, with per-year instruction data |
| Time Machine GPT | GPT-2 | 2011–2022 | yearly (12) | WMT News Crawl + Wikipedia, datasets released |
| TimeLMs | RoBERTa-base (and one large) | 2019-Q4–2022-Q4 | quarterly | |
| StoriesLM | BERT-style, expanding windows | 1900–1963 | yearly (64) | American Stories newspapers |
| HistBERT | BERT-base, continued pretraining | 1900s–2000s | decadal (10) | COHA |
Research threads
The literature clusters into a few threads. The literature map draws the connections; the timeline dates them.
Building the models
Train a sequence of checkpoints on chronologically filtered text. Recent work shows the performance gap to unconstrained models narrows with scale.
Measuring look-ahead bias
How much of an LLM's “forecasting skill” is memorised outcome? Detection statistics, fake-date tests, and finance-specific benchmarks.
Can prompting substitute?
Telling a model to “pretend it is 2015” is cheap. It works for direct queries and fails for causally downstream knowledge; models struggle with chronology itself.
Temporal generalisation
The NLP thread that preceded finance: models degrade as the world moves past their training window, and continual pretraining only partly helps.
Temporal knowledge benchmarks
What does a model know as of when? Cutoff tracing, chronological knowledge across domains, evolving medical guidelines.
Beyond finance
Historical LLMs as simulated informants for behavioural science, and proposed uses in medicine, law, history, journalism, and policy.
Corpora
Every checkpoint sequence is only as good as its dated text. The corpora page catalogues candidate sources by domain, with coverage, licensing, and the finest temporal granularity each supports, from Chronicling America and COHA to SEC EDGAR, Hansard, and Common Crawl derivatives such as FineWeb.
What to build next
The plans page holds three costed proposals for extending the open model and corpus sequences on a small budget: a focused $500k programme, an open-source corpus programme, and a decentralised Bittensor-subnet variant. They are working documents, not commitments.
Further reading
- He, Lv, Manela & Wu (2025). Chronologically Consistent Large Language Models — the paper that made the construction mainstream, with ChronoBERT and ChronoGPT.
- Kelly, Malamud, Schwab & Xu (2026). Scaling Point-in-Time Language Models — the gap to unconstrained models closes with scale.
- Ludwig, Mullainathan & Rambachan (2024). Large Language Models: An Applied Econometric Framework — why training leakage matters for inference, not just prediction.
- Point-in-Time LLMs Beyond Finance — the essay on this site: what temporally faithful models make possible in six other fields.
- Overview slides — a short presentation of the idea.
Bibliography
Grouped by thread. Every entry also appears on the literature map and the timeline, and has a one-paragraph note on the key ideas page. arXiv identifiers were checked against the arXiv API; SSRN and journal items against Crossref.
Model sequences
- He, S., Lv, L., Manela, A., and Wu, J. (2025). Chronologically Consistent Large Language Models. arXiv:2502.21206. ChronoBERT and ChronoGPT: 26 yearly checkpoints each, 1999–2024, trained only on text available at each date; competitive with BERT-class models on NLP benchmarks.
- He, S., Lv, L., Manela, A., and Wu, J. (2025). Instruction Tuning Chronologically Consistent Language Models. arXiv:2510.11677. A chat-style, instruction-tuned ChronoGPT per cut-off year, giving a conservative lower bound on forecast accuracy once leakage is removed.
- Kelly, B., Malamud, S., Schwab, J., and Xu, T. A. (2026). Scaling Point-in-Time Language Models. arXiv:2607.11889. Decoder-only models up to 4B parameters on 1T chronologically filtered FineWeb tokens; monthly checkpoints 2013–2024; the gap to open-weight frontier models narrows with scale.
- Drinkall, F., Rahimikia, E., Pierrehumbert, J. B., and Zohren, S. (2024). Time Machine GPT. Findings of NAACL 2024. arXiv:2404.18543. Twelve yearly “nonprognosticative” GPT-2 models (2011–2022) trained from scratch on WMT News Crawl and Wikipedia; models and training data on Hugging Face.
- Loureiro, D., Barbieri, F., Neves, L., Espinosa Anke, L., and Camacho-Collados, J. (2022). TimeLMs: Diachronic Language Models from Twitter. ACL 2022 demo. Quarterly RoBERTa checkpoints from 2019-Q4 to 2022-Q4, with a library that routes each tweet to the model that could have seen it.
- Sarkar, S. K. (2024). StoriesLM: A Family of Language Models With Time-Indexed Training Data. SSRN 4881024. Sixty-four yearly models, 1900–1963, each continuing the previous year's checkpoint on that year's American Stories newspaper text.
- Qiu, W., and Xu, Y. (2022). HistBERT: A Pre-trained Language Model for Diachronic Lexical Semantic Analysis. arXiv:2202.03612. Decade-wise BERT models continued on COHA, for tracking lexical semantic change.
- Röttger, P., and Pierrehumbert, J. B. (2021). Temporal Adaptation of BERT and Performance on Downstream Document Classification. Findings of EMNLP 2021. Early evidence on what temporal adaptation does and does not buy.
- Yan, Y., Tang, R., Gao, Z., and Jiang, W. (2026). DatedGPT: Preventing Lookahead Bias in Large Language Models with Time-Aware Pretraining. arXiv:2603.11838. Twelve 1.3B models with annual cut-offs 2013–2024 plus a per-year instruction set; prices the “lookahead premium” at 26.4 bp per standard deviation in return prediction.
- Faro, R., Fan, D., Alphaidze, T., and Jaggi, M. (2025). TiMoE: Time-Aware Mixture of Language Experts. arXiv:2508.08827. Experts pretrained on disjoint two-year slices; at inference, experts whose window ends after the query date are masked, so one set of weights serves every cut-off.
- Li, J., et al. (2025). TiC-LM: A Web-Scale Benchmark for Time-Continual LLM Pretraining. arXiv:2504.02107. 114 Common Crawl dumps; continual pretraining with replay matches retraining from scratch at 2.6× less compute.
- Pilchen, H., Fabre, R., Signe Talla, F., and Pérez, P. (2026). Understanding Data Temporality Impact on Large Language Models Pre-training. arXiv:2605.22769. Chronologically ordered pretraining yields fresher, more temporally precise knowledge than shuffled pretraining at no cost in general ability.
- Fittschen, E., Li, S., Lippincott, T., and Choshen, L. (2025). Pretraining Language Models for Diachronic Linguistic Change Discovery. Findings of EACL 2026. Five 10-million-word time slices; efficient pretraining beats fine-tuning an 8B model for period-faithful inference.
- Luo, X., Shinnick, Z., Griesshaber, N., and Wang, Y. (2026). Pretraining Language Models on Historical Text. arXiv:2606.02991. TypewriterLM, a 7B model on 54B tokens of pre-1913 English, with leakage-mitigated corpus construction and lexically grounded instruction tuning.
Look-ahead bias and training leakage
- Lopez-Lira, A., and Tang, Y. (2023). Can ChatGPT Forecast Stock Price Movements? arXiv:2304.07619. The result that prompted the leakage question: strong in-sample “forecasts” from a model that had read the outcomes.
- Glasserman, P., and Lin, C. (2023). Assessing Look-Ahead Bias in Stock Return Predictions Generated By GPT Sentiment Analysis. arXiv:2309.17322. Anonymising firms and dates to separate genuine sentiment signal from memorised outcomes.
- Sarkar, S. K., and Vafa, K. (2024). Lookahead Bias in Pretrained Language Models. SSRN 4754678. Direct measurement of leakage in pretrained models used for economic prediction.
- Rahimikia, E., and Drinkall, F. (2024). Re(Visiting) Large Language Models in Finance. SSRN 4963618. Financial applications revisited with point-in-time models in hand.
- Ludwig, J., Mullainathan, S., and Rambachan, A. (2024). Large Language Models: An Applied Econometric Framework. arXiv:2412.07031. When LLM outputs can and cannot be used as data in empirical research; training leakage is one of the two central problems.
- Lopez-Lira, A., Tang, Y., and Zhu, M. (2025). The Memorization Problem: Can We Trust LLMs' Economic Forecasts? arXiv:2504.14765. In-sample forecasting skill is non-identified; models recall exact values, and neither date instructions nor masking prevent it.
- Crane, D., Karra, A., and Soto, P. E. (2025). Total Recall? Evaluating the Macroeconomic Knowledge of Large Language Models. Federal Reserve FEDS 2025-044. LLMs smooth across data vintages and believe they hold data not yet released.
- Paleka, D., Goel, S., Geiping, J., and Tramèr, F. (2025). Pitfalls in Evaluating Language Model Forecasters. arXiv:2506.00723. A catalogue of temporal-leakage channels in forecaster evaluation.
- Li, C., et al. (2025). Time Travel is Cheating: Going Live with DeepFund for Real-Time Fund Investment Benchmarking. NeurIPS 2025. A live benchmark on post-cutoff market data; the only leak-proof backtest is a prospective one.
- Merchant, H., and Levy, B. (2025). A Fast and Effective Solution to the Problem of Look-ahead Bias in LLMs. arXiv:2512.06607. Logit steering of a large base model by two small models, one tuned on what to forget and one on what to keep.
- Levy, B. (2026). Caution Ahead: Numerical Reasoning and Look-Ahead Bias in AI Models. Journal of Accounting Research. Much of LLMs' apparent superhuman performance in accounting and finance is artifact: poor numerical reasoning plus look-ahead bias.
- Gao, Z., Jiang, W., and Yan, Y. (2025). Detecting Lookahead Bias in LLM Forecasts. arXiv:2512.23847. A “lookahead propensity” statistic from date-only recall queries; it collapses to zero after the training cutoff.
- Eliseev, A., and Seleznev, S. (2026). Fake Date Tests: Can We Trust In-sample Accuracy of LLMs in Macroeconomic Forecasting? arXiv:2601.07992. Prompt-sensitivity tests; none of the tested models passed.
- Benhenda, M. (2026). Look-Ahead-Bench: a Standardized Benchmark of Look-ahead Bias in Point-in-Time LLMs for Finance. arXiv:2601.13770. Alpha decay across market regimes as the measure; standard LLMs show it, point-in-time models do not.
- Kong, Y., Lee, H., Hwang, Y., Lopez-Lira, A., Levy, B., Mehta, D., Wen, Q., Choi, C., Lee, Y., and Zohren, S. (2026). Evaluating LLMs in Finance Requires Explicit Bias Consideration. arXiv:2602.14233. Position paper: five biases including look-ahead; a review of 164 papers finds none discussed in more than 28% of them.
- Zhang, Z., Chen, R., and Stadie, B. C. (2026). All Leaks Count, Some Count More: Interpretable Temporal Contamination Detection and Mitigation in LLM Backtesting. arXiv:2602.17234. Shapley-weighted, claim-level leakage rate; TimeSPEC grounds predictions in date-filtered evidence.
- Gao, Z., Jiang, W., and Yan, Y. (2026). Debiasing LLMs by Fine-tuning. arXiv:2604.02921. LoRA fine-tuning on rational benchmark forecasts fixes extrapolation bias where prompting could not.
- Zhang, Z., and Stadie, B. C. (2026). TEMPO: Temporal Enforcement via Mode-Separated Policy Optimization for Trustworthy LLM Backtesting. arXiv:2605.18843. Temporal compliance is instance-specific, so unlearning cannot work; train temporal discipline with a two-mode reward.
- Zhang, F., Li, Z., Peng, S., and Chen, Y. (2026). When Alpha Disappears: A One-Switch Benchmark for Decision-Time Leakage in Financial Backtests. arXiv:2605.23959. Toggle one evaluation convention at a time to measure protocol-induced inflation.
- Merchant, H., and Levy, B. (2026). Forecasting With LLMs: Improved Generalization Through Feature Steering. arXiv:2606.27199. Amplifying sparse-autoencoder features for time-awareness reduces look-ahead bias; steering the look-ahead features does nothing.
- Fonseca, X. (2026). Look-Ahead-Freedom as Temporal Non-Interference: A Verifiable Correctness Property for Backtesting and Agentic Trading Pipelines. arXiv:2607.04958. Look-ahead-freedom as an information-flow property, decidable in linear time on the value-independent fragment.
- Jia, H. (2026). HindsightBench: A Black-Box Behavioral Audit Protocol for Parametric Hindsight in Time-Indexed LLM Decision Tasks. arXiv:2607.18867. Probe-cost audit via a four-arm date-manipulation matrix; the date-trigger reflex tracks training recency, not scale.
- Zhang, Z., and Stadie, B. C. (2026). Temporal Leakage in LLM Backtesting: Measurement, Validation, and Adjusted Scores. arXiv:2608.02985. The pre/post-cutoff comparison is provably uninformative; measurement needs a known cutoff or a matched clean control.
- Li, W. W., Wang, M., and Ma, T. (2026). Summoning the Oracle to Slay It: Mitigating Look-Ahead Bias in Financial Backtesting with LLMs. arXiv:2605.24564. “Parametric look-ahead bias” and an inference-time decoding fix.
- Cai, X., et al. (2026). OpenPM: Auditable Point-in-Time Evaluation for LLM Portfolio-Management Agents. arXiv:2608.09988. Every record visible to the agent must be available at decision time; runs emit a contamination certificate.
Prompted cut-offs and chronology
- Gao, X., Zhang, R., Du, D., Mahindre, S., Somayajula, S. A., and Xie, P. (2025). Can Prompts Rewind Time for LLMs? Evaluating the Effectiveness of Prompted Knowledge Cutoffs. arXiv:2510.02340. Prompted cut-offs work for directly queried facts and fail for semantic shift and causally related knowledge.
- Asai, M., et al. (2026). Can LLMs Be Constrained to the Past? Improving Knowledge Cutoff through Recall-Based Prompting. arXiv:2606.05804. Recall-based prompting narrows, but does not close, the gap.
- Wongchamcharoen, P. K., and Glasserman, P. (2025). Do Large Language Models (LLMs) Understand Chronology? arXiv:2511.14214. Frontier models preserve local order but fail to maintain a globally consistent timeline, undermining prompt-based defences.
- Cheng, J., Marone, M., Weller, O., Lawrie, D., Khashabi, D., and Van Durme, B. (2024). Dated Data: Tracing Knowledge Cutoffs in Large Language Models. arXiv:2403.12958. The effective cutoff differs from the reported one, and differs by resource.
- Zhao, B., Brumbaugh, Z., Wang, Y., Hajishirzi, H., and Smith, N. A. (2024). Set the Clock: Temporal Alignment of Pretrained Language Models. arXiv:2402.16797. Aligning a model's internal “now” to a target year.
- Park, Y., Yoon, C., Park, J., Jeong, M., et al. (2025). Does Time Have Its Place? Temporal Heads: Where Language Models Recall Time-specific Information. ACL 2025. Specific attention heads carry time-specific knowledge and can be ablated or edited.
- Pęzik, P., Kaczyński, K., Szymańska, M., and Żarnecki, F. (2025). LLMLagBench: Identifying Temporal Training Boundaries in Large Language Models. arXiv:2511.12116. Infer undeclared training cutoffs from knowledge of recent events.
- Li, Z., Wang, Y., El Lahib, A., and Xia, Y.-J. (2026). Simulated Ignorance Fails: A Systematic Study of LLM Behaviors on Forecasting Problems Before Model Knowledge Cutoff. arXiv:2601.13717. A 52% gap between simulated and true ignorance across 477 questions and nine models.
- Ding, C., Wu, J., Luo, Y., Liu, Z., et al. (2026). Teaching Large Language Models When Not to Know: Learning Temporal Critique for Ex-Ante Reasoning. arXiv:2605.14636. Ex-ante correctness is a relation between answer and cutoff; train a critic rather than a memory wipe.
- van Adrichem, S., Bhaskar, A., Yang, D., and Potts, C. (2026). Do Language Models Consistently Encode the Current Year? arXiv:2608.15507. Associative and declarative probes for the current year use different mechanisms; nothing updates both.
Temporal generalisation and misalignment
- Lazaridou, A., et al. (2021). Mind the Gap: Assessing Temporal Generalization in Neural Language Models. NeurIPS 2021. Performance degrades on text from after the training period; the case for time-stratified evaluation.
- Luu, K., Khashabi, D., Gururangan, S., Mandyam, K., and Smith, N. A. (2021). Time Waits for No One! Analysis and Challenges of Temporal Misalignment. arXiv:2111.07408.
- Agarwal, O., and Nenkova, A. (2021). Temporal Effects on Pre-trained Models for Language Processing Tasks. arXiv:2111.12790.
- Dhingra, B., Cole, J. R., Eisenschlos, J. M., Gillick, D., Eisenstein, J., and Cohen, W. W. (2021). Time-Aware Language Models as Temporal Knowledge Bases. TACL. Prefix each training example with its date; the model learns to condition on time.
- Jang, J., et al. (2021). Towards Continual Knowledge Learning of Language Models. ICLR 2022.
- Jin, X., et al. (2021). Lifelong Pretraining: Continually Adapting Language Models to Emerging Corpora. arXiv:2110.08534.
- Zhang, M. J. Q., and Choi, E. (2023). Mitigating Temporal Misalignment by Discarding Outdated Facts. arXiv:2305.14824. Predict how long a fact stays true.
- Shin, C., et al. (2025). TARDIS: Mitigating Temporal Misalignment via Representation Steering. arXiv:2503.18693.
- Elbadry, R., Heakl, A., Zhang, F., Bouch, D., et al. (2026). The Geometry of Forgetting: Temporal Knowledge Drift as an Independent Axis in LLM Representations. arXiv:2605.09195. Staleness is encoded orthogonally to correctness and uncertainty, so uncertainty-based detectors miss it by construction.
Temporal knowledge benchmarks
- Park, Y., Yoon, C., Park, J., Lee, D., Jeong, M., and Kang, J. (2024). ChroKnowledge: Unveiling Chronological Knowledge of Language Models in Multiple Domains. ICLR 2025. ChroKnowBench spans general, biomedical, legal, commonsense, and mathematical knowledge, separating evolving from constant facts. Code.
- Jang, J., et al. (2022). TemporalWiki: A Lifelong Benchmark for Training and Evaluating Ever-Evolving Language Models. EMNLP 2022. Code.
- Liška, A., et al. (2022). StreamingQA: A Benchmark for Adaptation to New Knowledge over Time in Question Answering Models. ICML 2022.
- Kasai, J., et al. (2022). RealTime QA: What's the Answer Right Now? arXiv:2207.13332.
- Deng, H., Jiao, W., Liu, X., Zhang, M., and Tu, Z. (2024). NewTerm: Benchmarking Real-Time New Terms for Large Language Models with Annual Updates. arXiv:2410.20814.
- Guan, Z., et al. (2026). Large Language Models Lack Temporal Awareness of Medical Knowledge. arXiv:2605.13045. TempoMed-Bench: evolving clinical guidelines as the test of “knowing when knowledge is correct”.
- Vladika, J., Dhaini, M., and Matthes, F. (2025). Facts Fade Fast: Evaluating Memorization of Outdated Medical Knowledge in Large Language Models. EMNLP 2025. MedChangeQA: 512 questions where medical consensus changed; all models tested answer with the old consensus.
- Xian, R. P., Cui, Q., Bauer, S., and Abbasi-Asl, R. (2025). Measuring temporal effects of agent knowledge by date-controlled tool use. REALM 2025. Date-restricted search tools as a stress test for agent temporal knowledge.
- Yuan, Z., Ding, Z., and Vlachos, A. (2025). Do Language Models Update their Forecasts with New Information? arXiv:2509.23936. EvolveCast: models move forecasts in the right direction but far too little given post-cutoff evidence.
- Zheng, J., Shao, Z., Ritter, A., and Xu, W. (2026). Synthetic Worlds for Temporal Evaluation and Knowledge Updating in LLMs. arXiv:2609.00184. Fictional but coherent future worlds give contamination-free temporal evaluation.
Corpora and data
- Dell, M., et al. (2023). American Stories: A Large-Scale Structured Text Dataset of Historical U.S. Newspapers. NeurIPS 2023 Datasets. Article-level text from nearly 20 million Chronicling America scans; the training data for StoriesLM.
- Silcock, E., Arora, A., D'Amico-Wong, L., and Dell, M. (2024). Newswire: A Large-Scale Structured Database of a Century of Historical News. NeurIPS 2024 Datasets. 2.7 million unique public-domain newswire articles, 1878–1977, reconstructed by deduplicating 138 million local-newspaper articles.
- Yang, H., Liu, X.-Y., and Wang, C. D. (2023). FinGPT: Open-Source Financial Large Language Models. arXiv:2306.06031. Data-centric financial LLM pipeline; not point-in-time, but a source of dated financial text.
- See the corpora page for the full catalogue of dated text sources.
Beyond finance
- Varnum, M. E. W., Baumard, N., Atari, M., and Gray, K. (2024). Large Language Models based on historical text could offer informative tools for behavioral science. PNAS 121(42). Historical LLMs as simulated informants for the psychology of past populations.
- Point-in-Time LLMs Beyond Finance. This site. Medicine, law, history, journalism, policy, and the social sciences: what becomes testable when the model cannot see ahead.
Working on a point-in-time model, corpus, or benchmark that is missing here? Open an issue or pull request on microprediction/pitllm.