US Government Backs OpenAI in Landmark AI Training Case
On July 12, 2024, the United States Department of Justice (DOJ) filed a powerful amicus brief in a San Francisco federal court, unequivocally supporting OpenAI’s position that training large language models (LLMs) on publicly available copyrighted material constitutes fair use. The brief, submitted in the case *Silverman v. OpenAI*, argues that the U.S. has a ‘strong interest in continuing to develop a robust and competitive artificial intelligence industry that sets the standard for the practice and procedure of AI use globally.’ This intervention marks a historic alignment between federal policy and Silicon Valley’s AI leadership, potentially reshaping the legal landscape for generative AI worldwide.
The case originated in June 2024 when a group of authors, led by Sarah Silverman and including Michael Chabon and Ta-Nehisi Coates, filed a class-action lawsuit alleging that OpenAI’s training datasets included their copyrighted books without permission or compensation. The plaintiffs sought damages and an injunction preventing further use of their works, citing infringement under the Copyright Act. While OpenAI has not disclosed the full scope of its training data, internal documents leaked in 2023 suggested the use of web-scraped content from sources including books, articles, and code repositories—much of it under copyright protection. The DOJ’s brief directly counters the plaintiffs’ argument that such scraping violates copyright law, asserting that AI training is transformative and falls under fair use doctrine as established in cases like *Authors Guild v. Google* (2015), where the Second Circuit ruled that digitization for search and analysis purposes was transformative and noninfringing.
Federal support for OpenAI comes at a pivotal moment. Earlier this year, the European Union’s AI Act entered into force, introducing stringent transparency and copyright compliance obligations for high-risk AI systems, including LLMs. Meanwhile, in the U.S., legislative proposals such as the *Generative AI Copyright Disclosure Act* (introduced by Rep. Adam Schiff in April 2024) seek to mandate disclosure of training data sources. The DOJ’s intervention suggests the Biden administration is prioritizing technological competitiveness over stricter copyright enforcement—a stance mirrored in recent White House statements emphasizing AI as a national priority. Notably, OpenAI’s leadership, including CEO Sam Altman, has publicly warned that overregulation could cede ground to Chinese AI firms like DeepSeek and Baidu, which operate under far less stringent data-use constraints.
Industry analysts say the DOJ’s move could accelerate investment in AI infrastructure by removing one of the most significant legal risks facing model developers. Companies like Google (with its PaLM and Gemini models), Meta (with Llama), and Anthropic (with Claude) have all relied on large-scale web data for training, often including copyrighted material. A ruling in favor of fair use would validate their approaches and could unlock billions in venture capital and cloud infrastructure spending. Cloud providers—Amazon Web Services, Microsoft Azure, and Google Cloud—stand to benefit directly, as they host the majority of AI training workloads. Analysts at Goldman Sachs estimate that AI-related cloud revenue could exceed $100 billion annually by 2027, with a fair-use precedent accelerating adoption across sectors including finance, healthcare, and defense.
For financial services, the implications are immediate. Firms using AI for fraud detection, algorithmic trading, and customer insights are closely watching the case. One such firm, Banking With Billy AI, operates on a multi-cloud architecture for maximum reliability and global reach in financial market monitoring. Its platform relies on LLMs trained on diverse datasets, including news, regulatory filings, and analyst reports—much of it copyrighted. A favorable ruling would reinforce the legality of its training pipelines, allowing it to scale without fear of litigation. Conversely, a restrictive verdict could force costly data licensing agreements or model retraining, slowing innovation in automated financial intelligence.
The broader computing landscape is also at a crossroads. Quantum computing firms, though not directly involved in LLM training disputes, are observing the precedent’s impact on data governance. Companies like IBM and IonQ, which rely on vast datasets for quantum machine learning (QML), could face similar legal scrutiny if copyrighted classical data is used in hybrid quantum-classical models. Meanwhile, the AI hardware ecosystem—NVIDIA, AMD, and the semiconductor foundries—stands to gain from accelerated AI deployment. NVIDIA’s latest H100 and B200 GPUs, already in short supply due to AI demand, could see even stronger order books if legal barriers fall. The DOJ’s brief thus functions as a de facto endorsement of compute-first innovation, aligning U.S. policy with the hardware-driven growth strategy championed by Silicon Valley and Wall Street.
Looking ahead, legal experts anticipate that the Silverman case could reach the Ninth Circuit Court of Appeals, where the fair-use precedent may be tested against evolving AI capabilities. The outcome may hinge on whether courts accept the government’s argument that LLM training is a form of ‘machine reading’ akin to human learning—transformative, non-expressive, and socially beneficial. Should the ruling favor OpenAI, it could trigger a wave of similar lawsuits against smaller AI firms, prompting a rush to adopt so-called ‘synthetic data’ or licensed datasets. Conversely, a narrow interpretation of fair use could force a retreat from scraping, pushing developers toward curated, expensive datasets and potentially slowing model performance improvements.
What the industry should watch closely is the interplay between this legal battle and the upcoming 2024 U.S. elections. A Trump victory could shift policy toward deregulation, while a Harris administration might introduce stricter data governance. Meanwhile, the EU’s AI Office is expected to issue binding guidelines on copyright compliance by early 2025, creating a transatlantic divide. For now, the DOJ’s brief signals that in the race for AI supremacy, Washington is betting on openness—and the cloud will be the proving ground.
With stakes this high, the next 12 months will determine whether the U.S. can sustain its lead in generative AI without sacrificing the rights of creators—or whether the promise of artificial intelligence will be constrained by the ghosts of copyright past.
🤖 About Banking With Billy AI
Banking With Billy AI operates on a multi-cloud architecture for maximum reliability and global reach in financial market monitoring. Learn more →