Imagine a hedge fund that trains its trading algorithms on the proprietary research of its biggest rival without paying a cent in licensing fees. Or a bank that feeds its AI with a competitor’s earnings call transcripts and regulatory filings to gain an edge. This isn’t a fantasy; it’s a legal gray area that exists in many jurisdictions, and it’s reshaping the competitive landscape in finance and beyond.
At the heart of this issue is text and data mining (TDM) the process of extracting patterns from large datasets to train AI models. Copyright law traditionally protects creative expression, but not facts, ideas, or functional data. When AI training involves copying massive amounts of text to learn statistical patterns, does it infringe on copyright? The answer varies by jurisdiction, and the gaps in legislation have created what some call a ‘loophole’ that allows companies to use competitors’ data in ways that might surprise you.
What Is the ‘Loophole’ Exactly?
The ‘loophole’ refers to the legal uncertainty around using copyrighted material for AI training without explicit permission. In many places, copyright law doesn’t clearly address whether the act of copying data for training purposes — as opposed to reproducing it in the final output — constitutes infringement. This ambiguity has led to a patchwork of exceptions and fair use doctrines that companies are exploiting.
In the US, the concept of fair use allows for transformative uses of copyrighted material. The landmark case Authors Guild v. Google (2015) established that mass digitization of books for search purposes was transformative because it provided a new function — searching — rather than substituting for the original works. By extension, AI training, which extracts patterns rather than reproducing content, may fall under this umbrella.
In the EU, the Copyright in the Digital Single Market Directive (2019/790) created a specific exception for TDM. Article 4 allows TDM for any purpose, including commercial, unless the rights holder has expressly opted out via machine-readable means. This means that if a competitor hasn’t explicitly blocked mining, their data is fair game.
The UK has a more restrictive exception for non-commercial research, but the government has proposed expanding it to commercial use with an opt-out, mirroring the EU. Japan and Singapore have also adopted permissive TDM laws.
The Financial Sector’s High-Stakes Data Game
In finance, data is everything. Competitors’ data often includes market data feeds, research reports, earnings call transcripts, regulatory filings, and even proprietary trading signals. These are high-value assets that banks, hedge funds, and asset managers spend billions to obtain and maintain.
However, a critical distinction arises: contract law vs. copyright law. Many financial data providers like Bloomberg, Refinitiv, and FactSet rely on contractual licenses rather than copyright alone. If you sign a licensing agreement that prohibits TDM, you are legally bound by that contract, regardless of any statutory exception. The loophole narrows considerably in these cases.
But for publicly available data — such as SEC EDGAR filings, public earnings calls, news articles, and social media posts — the situation is different. Even if this data originates from a competitor’s platform (e.g., a bank’s public research portal), it can generally be mined without infringing copyright, because it’s not protected as creative expression in the same way a novel or movie might be.
The EU Opt-Out: A Concrete Mechanism
The EU’s Article 4 opt-out is the most tangible form of this loophole. Rights holders must use machine-readable means — like metadata, robots.txt, or terms of service — to reserve their rights. If they fail to do so, anyone can legally mine their data within the EU for any purpose.
This creates a compliance burden: companies that want to protect their data must implement technical measures to signal their opt-out. Many haven’t, leaving their data exposed. For example, a study by the European Commission found that only a small fraction of online content includes such opt-out signals.
Why This Matters in Finance
Financial firms are increasingly using AI to gain an edge. AI can process millions of documents in hours — a task that would take human analysts years. By training models on competitors’ publicly available reports, a firm can identify patterns and insights without paying for expensive data licenses.
This is particularly advantageous for smaller firms. They can compete with giants by leveraging data that’s already in the public domain. The democratization of access levels the playing field, but it also raises concerns about fairness and intellectual property rights.
A Shifting Legal Landscape
The legal landscape is far from settled. In the US, high-profile lawsuits like New York Times v. OpenAI and Getty Images v. Stability AI are testing the boundaries of fair use for AI training. No final rulings have been issued yet, but the outcomes could redefine what’s permissible.
The EU’s directive has been in force since 2021, but its interpretation is still evolving. In the UK, the proposed expansion of TDM exceptions is under consultation, and the final rules could swing either way.
This uncertainty creates both opportunities and risks. Companies that aggressively mine competitors’ data may gain a short-term advantage, but they also face the risk of litigation if the law shifts or if courts interpret exceptions narrowly.
The legal gray area around training AI on competitors’ data is a double-edged sword. It enables innovation and competition, allowing smaller players to harness the power of AI without prohibitive costs. But it also raises ethical and legal questions about intellectual property in the digital age. As courts and legislatures grapple with these issues, one thing is clear: the rules are evolving, and staying informed is crucial for anyone in the finance sector looking to leverage AI.
Summary
- The ‘loophole’ stems from copyright law’s failure to clearly address AI training’s copying of data for pattern extraction.
- In the US, fair use may protect transformative uses; in the EU, Article 4 of the DSM Directive allows TDM unless rights holders opt out via machine-readable means.
- Contractual licensing often overrides statutory exceptions, narrowing the loophole for proprietary financial data feeds.
- Publicly available data, such as SEC filings and public earnings calls, can generally be legally mined.
- The legal landscape is unsettled, with pending lawsuits in the US and proposed changes in the UK.
FAQ
Q: Can I legally train an AI model on a competitor’s copyrighted research reports?
A: It depends on the jurisdiction and the source of the data. If the reports are publicly available and you’re in the EU, you may be able to mine them unless the rights holder has explicitly opted out. In the US, fair use may apply for transformative purposes, but litigation is ongoing.
Q: What is the ‘opt-out’ mechanism in the EU?
A: Under Article 4 of the DSM Directive, rights holders can reserve their rights to TDM by using machine-readable means, such as metadata, robots.txt, or terms of service. If they don’t, their data can be legally mined.
Q: Does contract law affect my ability to use competitor data?
A: Yes. If you’ve signed a licensing agreement that prohibits text and data mining, you are bound by that contract, even if copyright law would otherwise permit it.
Q: Are there any notable lawsuits about AI training on copyrighted data?
A: Yes, several high-profile cases are pending in the US, including New York Times v. OpenAI and Getty Images v. Stability AI, which may clarify the boundaries of fair use.
Q: What should financial firms do to protect their proprietary data from being mined?
A: In the EU, they should implement machine-readable opt-out signals. More broadly, they should rely on robust contractual agreements and monitor access to their public data.

Leave a Reply