This piece first appeared in German. The English version is a rewrite rather than a line-by-line translation, and the German original stays online in the archive: Wie maschinelles Lernen die Finanzindustrie seit Jahrzehnten prägt.


Banks were quantifying credit risk with mathematical models in the 1980s. Income, employment status, payment history went into an equation and a probability of default came out. The model did not learn anything; it calculated. That was the first move away from a purely subjective lending decision toward a reproducible, scalable one.

The finance industry is returning to artificial intelligence, and not discovering it. Risk departments, trading desks and lenders have worked with learning systems for decades, and the difference between a credit scoring model from the 1990s and a modern language model is real without being fundamental. It is a difference of degree.


Explainability was the original selling point

Those early models had one obvious strength. Every input could be weighted and explained, so a regulator could reconstruct why an applicant was turned down. That property becomes one of the central regulatory disputes later, in the era of models with billions of parameters.

The 1990s brought methods that found patterns people could no longer see directly. A decisive break with classical statistics came with it: these models were not programmed with fixed rules. They learned their rules from example data, so a fraud detection system no longer had to be told what fraud looks like: it learned from thousands of documented cases. Conceptually that was new, even where the hardware of the time limited how complex such a model could get.


Credit scoring is the oldest scalable machine learning field

Models assess an applicant's creditworthiness automatically by recognizing patterns in historical lending data, and the output is a number that banks and card issuers use as a basis for decisions.

What began in the 1980s as a simple statistical equation grew into more complex learning systems over decades. One advantage stayed constant: trained once, these models evaluate millions of applications consistently and without fatigue. The challenge stayed constant too. Historical data reflects historical inequality, and a model trained on the past can disadvantage particular groups systematically, which regulation has still not fully solved.


Algorithmic trading moved from rulebook to learned pattern

Learning systems have been identifying price patterns and reacting to them automatically since the early 2000s. The first systems were rule-based: if price X falls below threshold Y, buy. People wrote the rules and computers executed them.

Machine learning changed that logic at the root. Models are no longer given rules; they learn for themselves which patterns in historical market data carry predictive power. Lopez de Prado documents the path from manually programmed rulebooks to systems that recalibrate market patterns continuously, faster and more consistently than any human trader. Market structure changed with it, and substantial shares of volume on major exchanges are now executed by algorithmic systems.

The downside is familiar. These systems can amplify instability. In the flash crash of 6 May 2010 the Dow Jones fell almost 1,000 points within minutes and recovered just as fast, driven in large part by the interaction of automated trading systems.


Text analysis in finance is older than the language models

Analysts wanted to know early what company reports actually say, beyond the official figures, and automated text analysis looked like an answer.

The tools available then had been built for general language and not for financial text. Loughran and McDonald showed that generic approaches fail systematically on filings. Take the word liability, which a general sentiment dictionary counts as negative and which in a filing with the Securities and Exchange Commission is usually a neutral accounting term. A finance-specific lexicon, which they built, laid the methodological groundwork for automated analysis of company reports and earnings calls.

The consequence reaches well past that one paper. Domain specificity is a precondition for reliable results and not a luxury. Apply a general model to financial text and the output is systematically distorted, however capable that model is elsewhere.


The jump to modern AI

In 2017 one publication changed direction. Vaswani and colleagues introduced an architecture that replaced the sequential processing of language that had dominated until then.

What is a transformer? Older models processed text word by word, the way a person reads a sentence from left to right. A transformer analyzes all the words at once and evaluates which of them matter to each other. Bank means one thing in "I am sitting on the bank" and another in "the bank issues loans", and a transformer distinguishes them from the surrounding context. That makes long-range relationships visible and makes training far faster and more scalable.

What followed was a bet on scale: more parameters, more training data, more compute, and capabilities that smaller models structurally do not show. An engineering optimization turned out to be a qualitative jump. Past a certain size, abilities appear that nobody trained explicitly, among them translation, logical inference and code generation.

For finance that became concrete in 2023 with BloombergGPT, a language model trained on 363 billion tokens of financial text from Bloomberg's archive, covering web pages, news, filings and press releases from 2007 to 2022, alongside 345 billion tokens of general text. It outperforms generic models on finance-specific tasks by a clear margin, which confirms Loughran and McDonald's finding at a new technological level. Domain specificity decides quality.

BloombergGPT is not a product anyone uses directly. It is a research result showing what becomes possible when training data and architecture are aimed consistently at one industry. The practical consequences run from automated news evaluation through real-time sentiment analysis to structured extraction from regulatory documents.

Two techniques established themselves alongside it. Fine-tuning adapts an already pretrained model to specific tasks with a limited, domain-specific dataset, much as a generalist goes through a specialization. Retrieval-augmented generation takes another route: the model reaches out to external, current knowledge sources at runtime and holds less internally. For compliance and research work that matters, because the sources stay traceable and the model does not operate on stale training data.


What changed, and what did not

The line of development is coherent, and the qualitative jump brings structurally new problems.

Older risk models are explainable. You can reconstruct which inputs produced which result and justify that path to a regulator. Modern language models with billions of parameters are structurally not explainable in that way, and the internal logic behind a particular output cannot be fully reconstructed even by the people who built it.

That is not an academic problem. The EU AI Act classifies AI systems in regulated decision processes such as lending and risk assessment as high-risk, with explicit requirements for transparency, traceability and human oversight. Basel IV raises the bar on model validation and documentation in risk management at the same time. Anyone integrating modern language models into those processes has to meet requirements designed for interpretable models, using systems that work on a different principle.

Hallucination compounds it. Language models generate plausible-sounding statements that are factually wrong, not because the model lies, and because it produces statistically likely continuations without checking whether they are true. In a research tool that is irritating and correctable. Put it in a risk model whose output feeds decisions directly and it becomes a systemic problem.

The industry has decades of practice in interrogating models, validating them and justifying them to regulators. Quantitative analysts have always known that no model represents reality completely, and that the real question is which simplifications are acceptable. That institutional knowledge is not an obstacle to adopting modern AI. It is the precondition for using it sensibly.


Sources

Grouped by the section they support.

Opening

  • myFICO (2018), The History of the FICO Score, 21 August 2018, myfico.com. Supports Fair Isaac, founded in 1956, and the introduction of the FICO Score in 1989. The description of early scoring models as equations is the author's summary.

Explainability was the original selling point

  • Ghosh and Reilly (1994), Credit Card Fraud Detection with a Neural-Network, Proceedings of the 27th Hawaii International Conference on System Sciences, ieeexplore.ieee.org. Supports a fraud detection system trained on documented fraud cases, here on Mellon Bank transaction data from 1991. The point about explainability to regulators is the author's own.

Credit scoring is the oldest scalable machine learning field

  • myFICO (2018). Supports scores as a basis for lending decisions. The points on consistency at scale and on historical inequality in training data are the author's own.

Algorithmic trading moved from rulebook to learned pattern

  • Lopez de Prado (2018), Advances in Financial Machine Learning, Wiley, chapter 1, Financial Machine Learning as a Distinct Subject, wiley-vch.de. Cited by the author for the move from programmed rules to learned models.
  • Staffs of the CFTC and SEC (2010), Findings Regarding the Market Events of May 6, 2010, 30 September 2010, sec.gov. Supports the flash crash, in which major equity indices fell a further 5 to 6 percent within minutes before recovering, and the finding that the interaction of automated execution programs and algorithmic trading strategies can quickly erode liquidity.

Text analysis in finance is older than the language models

  • Loughran and McDonald (2011), When Is a Liability Not a Liability? Textual Analysis, Dictionaries, and 10-Ks, Journal of Finance 66 (1), 35–65, doi.org/10.1111/j.1540-6261.2010.01625.x. Supports the failure of general-purpose word lists on company filings, with liability, capital and mine as examples, and the finance-specific word lists the authors built.

The jump to modern AI

  • Vaswani, Shazeer, Parmar, Uszkoreit, Jones, Gomez, Kaiser and Polosukhin (2017), Attention Is All You Need, NeurIPS 2017, arXiv:1706.03762. Supports the transformer and its departure from sequential processing. The bank example in the box is an illustration.
  • Wei et al. (2022), Emergent Abilities of Large Language Models, Transactions on Machine Learning Research, arXiv:2206.07682. Supports abilities that appear in larger models and not in smaller ones. Schaeffer, Miranda and Koyejo (2023), Are Emergent Abilities of Large Language Models a Mirage?, NeurIPS 2023, arXiv:2304.15004, argue that the effect depends on the metric chosen.
  • Wu, Irsoy, Lu, Dabravolski, Dredze, Gehrmann, Kambadur, Rosenberg and Mann (2023), BloombergGPT: A Large Language Model for Finance, arXiv:2303.17564. Supports the 363-billion-token financial dataset, its composition, and the lead over existing models on financial tasks.
  • Lewis et al. (2020), Retrieval-Augmented Generation for Knowledge-Intensive NLP Tasks, NeurIPS 2020, arXiv:2005.11401. Supports retrieval-augmented generation and its advantage for provenance and current knowledge. Fine-tuning is a standard term and needs no source.

What changed, and what did not

  • Regulation (EU) 2024/1689 (AI Act), Annex III point 5 and Articles 13 (transparency) and 14 (human oversight), AI Act Service Desk. Supports the high-risk classification of creditworthiness assessment and the related requirements.
  • Basel Committee on Banking Supervision (2017), Basel III: Finalising post-crisis reforms, 7 December 2017, bis.org. Supports the tighter constraints on banks' internal models, including the output floor.
  • Ji et al. (2023), Survey of Hallucination in Natural Language Generation, ACM Computing Surveys 55 (12), arXiv:2202.03629. Supports hallucination as fluent but false output. The closing argument about institutional knowledge is the author's own.
The link has been copied!