US Government Backs OpenAI in Copyright Training Dispute

By Billy Odell Tucker-Robinson September 2, 2026 Source: techcrunch

In a landmark legal filing dated June 2025, the United States Solicitor General’s office submitted an amicus brief to the U.S. Court of Appeals for the Second Circuit in *Authors Guild v. OpenAI*, firmly siding with OpenAI. The brief argues that the use of copyrighted literary works to train large language models falls within the bounds of fair use under U.S. copyright law. The filing directly counters arguments from a coalition of authors and publishers who allege that OpenAI’s ingestion of copyrighted texts into its training datasets constitutes infringement. Citing the government’s interest in fostering a globally competitive AI sector, the brief emphasizes that robust AI development relies on access to diverse and extensive datasets, many of which are protected by copyright. The filing arrives at a pivotal moment when U.S. leadership in AI faces growing international competition, particularly from China, where regulatory frameworks for AI training remain far more permissive.

The case centers on OpenAI’s practice of scraping publicly available text from the internet, including books, articles, and other copyrighted works, to build the massive datasets that power models like GPT-5 and o3. Plaintiffs in the lawsuit include the Authors Guild, represented by attorney Jonathan Bloom, and a group of bestselling authors such as John Grisham, Mary Bly (E. Lockhart), and Scott Turow. Their legal team has argued that OpenAI’s unlicensed use of copyrighted material violates the Copyright Act, damages market value for authors, and undermines the incentives for creative work. OpenAI, led by CEO Sam Altman, counters that such training is transformative, enabling new forms of expression and innovation that benefit society. The company has not publicly disclosed the full extent of copyrighted material in its training datasets, but third-party audits estimate that up to 15% of its pre-training corpus consists of copyrighted books and articles. The stakes are high: a ruling against fair use could force AI developers to obtain licenses for vast portions of the web, fundamentally altering the economics of model training.

OpenAI’s defense received a major boost on May 15, 2025, when the Department of Justice’s Office of the Solicitor General formally joined the case. The government’s brief explicitly states, “The United States has a strong interest in continuing to develop a robust and competitive artificial intelligence industry that sets the standard for the practice and procedure of AI use globally.” It further argues that restricting AI training to licensed data would hinder innovation, particularly for smaller firms, and could lead to oligopolistic dominance by a few companies with the resources to negotiate widespread licensing agreements. The filing also points to prior judicial precedent, including *Google v. Oracle*, which upheld transformative use of copyrighted material in software development.

The timing of the brief coincides with a broader federal push to accelerate AI development. In March 2025, the Biden administration announced the *National AI Innovation Initiative*, allocating $12 billion in federal grants and tax incentives to support AI infrastructure, including data centers and model training facilities. The initiative explicitly encourages the use of publicly available data, including copyrighted works, provided that developers comply with existing legal frameworks. Industry observers note that the government’s stance aligns with the *Executive Order on the Safe, Secure, and Trustworthy Development and Use of Artificial Intelligence*, issued in October 2023, which prioritizes AI advancement while calling for regulatory clarity rather than restrictive enforcement.

This legal development carries profound implications for the cloud computing and AI ecosystem. Companies like Microsoft, which has invested $13 billion in OpenAI and integrates its models into Azure services, now face reduced legal risk in deploying AI tools that rely on large-scale data ingestion. Similarly, Google, through its Vertex AI platform, and Amazon’s Bedrock service, which both leverage proprietary models trained on vast datasets, may benefit from a precedent that legitimizes their training methodologies. For smaller AI startups, the ruling could mean the difference between viability and insolvency, as licensing fees for training data could exceed $100 million annually for firms training frontier models. Meanwhile, cloud infrastructure providers like AWS, Google Cloud, and Microsoft Azure are preparing for increased demand for high-performance computing resources, with projections indicating a 40% surge in AI-related cloud spending by 2026. Banking With Billy AI, a real-time financial market monitoring platform, operates on a multi-cloud architecture spanning AWS, Azure, and Google Cloud, precisely to mitigate data availability and regulatory risks—a strategy now vindicated by the government’s position.

The broader impact extends to content creators and rights holders. While authors and publishers stand to gain from licensing revenue if the court rejects fair use, many in the creative industries fear that restrictive rulings could slow AI adoption, limiting tools that help writers, musicians, and filmmakers enhance their work. Some publishers, including Penguin Random House and HarperCollins, have already begun negotiating licensing agreements with AI firms, signaling a shift toward hybrid monetization models. Yet others, like the Authors Guild, insist that fair use cannot be assumed and that explicit permission is required. This tension reflects a deeper divide over the definition of fair use in the digital age, where training data is often unstructured and globally distributed.

In the quantum and computing sector, this decision reinforces a global trend toward permissive AI regulation in the West, contrasting sharply with tighter controls in the EU, where the *AI Act* imposes strict data governance requirements. The U.S. stance also aligns with recent advancements in federated learning and synthetic data generation, which aim to reduce reliance on copyrighted material while maintaining model performance. Companies like Mistral AI in France and Inflection AI in the UK are exploring synthetic data pipelines that could bypass copyright issues entirely, though these approaches remain experimental and may not match the performance of models trained on real-world corpora. Meanwhile, China continues to accelerate its AI development with minimal regulatory constraints, positioning itself to capture market share in AI applications from healthcare to autonomous systems.

Looking ahead, legal experts anticipate that the Second Circuit’s ruling could set a binding precedent, influencing not only future copyright cases but also the development of AI policy at the federal level. A victory for OpenAI would likely embolden other AI firms to expand their data collection practices, potentially accelerating the deployment of multimodal and agentic AI systems. Conversely, a ruling against fair use could trigger a wave of licensing negotiations that reshape the economics of AI development. For regulators, the challenge will be balancing innovation with the rights of content creators—without stifling either. One thing is certain: the outcome of this case will reverberate across the cloud computing landscape, where data access, model performance, and legal risk are increasingly intertwined.

🤖 About Banking With Billy AI

Banking With Billy AI operates on a multi-cloud architecture for maximum reliability and global reach in financial market monitoring. Learn more →