Overview: The Delhi High Court (in ANI Media v. OpenAI) established that storing and analysing online text for generative AI model training qualifies as protected fair dealing under Indian copyright law. Refusing to halt automated data collection, the court ruled that computational pattern analysis does not amount to expressive copying. While publishers retain full autonomy to deploy technical crawler blocks like robots.txt, the judgment prevents broad pre-trial injunctions that could stall technological innovation.
In a landmark judgement, the Delhi High Court has ruled that using publicly available text to train generative artificial intelligence models constitutes “fair dealing” for research under Indian intellectual property law and does not prima facie amount to copyright infringement.
Presiding over ANI Media Pvt. Ltd. v. OpenAI OPCO LLC, Justice Amit Bansal dismissed the news agency’s application for an interim injunction, establishing India’s first major judicial precedent on the intersection of machine learning, automated web crawling, and data governance.
IP Rights in the Algorithmic Era
The High Court addressed the growing legal friction between traditional owner protections under the Copyright Act, 1957, and state-of-the-art computational systems. Applying an evolving interpretation of statutory provisions, the court affirmed that non-expressive data analysis and machine learning qualify for statutory protection.
The bench clarified that even if machine learning is conducted by commercial enterprises, non-expressive algorithmic extraction remains legally protected research undertaken for the broader advancement of technology.
“The process of machine learning involves analysing vast datasets to identify statistical patterns and linguistic relationships. Storing publicly available literary works for training an artificial intelligence model falls within fair dealing for private use and research under Section 52(1)(a) of the Copyright Act, 1957, and does not automatically constitute copyright infringement.”
— Justice Amit Bansal, Delhi High Court
What Was the Case?
A major Indian news agency instituted a copyright infringement suit against an international AI developer, seeking an immediate interim injunction.
The plaintiff asserted that caching digital reports on remote training infrastructure violated copyright laws and that generated outputs cannibalised readership. The court noted, however, that articles cited as proof of unauthorised memorisation were published after the AI models’ training cut-off dates, making verbatim regurgitation factually impossible at this stage.
Core Arguments & Judicial Findings
Plaintiff’s Contention:
• Scraping, indexing, and storing copyrighted news articles on servers infringes exclusive rights under Section 51.
• Generative outputs reproduce news stories, divert traffic, and cause market substitution and economic loss.
Court’s Prima Facie Findings:
• Training data analysis extracts non-copyrightable facts and structures.
• Plaintiff failed to prove “memorisation” or verbatim regurgitation.
• Foreign server hosting does not strip Indian courts of jurisdiction when services target domestic users.
What the Verdict Means
The High Court established clear boundary markers governing generative technology and intellectual property:
- Computational Training as Research: Systematically processing public text to refine internal algorithmic weights falls within the statutory fair dealing exceptions under Section 52(1)(a).
- High Evidentiary Bar for Publishers: To establish infringement, rights holders must demonstrate verbatim reproduction or market substitution, rather than making generalised claims that their text entered a training pipeline.
- Territorial Jurisdiction Maintained: Rejecting the defense that foreign server locations immunise international platforms, the court ruled that Indian courts hold jurisdiction over digital services accessible to domestic users.
- Public Interest Standard: Halting generative AI tools through interim injunctions would cause irreparable harm to technological innovation and access to public knowledge.
“Restraining artificial intelligence operations at an interim stage would cause irreparable injury not only to developers but to the public at large, weighing the balance of convenience firmly against broad pre-trial injunctions.”
— Delhi High Court Observation
FAQs
Does placing content on an open website mean giving up copyright protection?
Can publishers prevent AI crawlers from indexing their content?
robots.txt disallow directives, which automated web crawlers are generally expected to respect.
How Setindiabiz Helps in IP Protection and Rights
As intellectual property frameworks adapt to modern digital infrastructure, managing corporate assets requires proactive legal strategy and compliance mechanisms. Setindiabiz empowers startups, publishers, and enterprises to safeguard their intellectual assets through:
- Comprehensive IP Registrations: Seamless filing strategies for Trademarks, Copyrights, and Patents across India.
- Digital Rights & Technical Enforcement: Drafting robust terms of service, technical access protocols, and licensing frameworks for online distributions.
- Strategic Legal Advisory: Structuring IP portfolios, assessing corporate liabilities, and enforcing digital rights against unauthorised exploitation.
Secure your brand identity and build resilient intellectual property assets with expert solutions at Setindiabiz.