
India's DPDP Act and Your AI: What Operators Need to Know Now
You have a customer database of 500K IDs, behavioral events, purchase history, and browsing patterns. You want to train a recommendation model. Under India's Digital Personal Data Protection Act (DPDP), passed in August 2023 and now in phased enforcement, you cannot. Not unless you rewired your consent flows months ago and documented a lawful basis for each data point you feed the model.
This is not a regulatory edge case. EY's DPDP compliance survey found that 81% of Indian companies have not updated their privacy policies or governance frameworks, and 83% have not begun comprehensive implementation. If your revenue spans India, or your AI ops touch Indian users, the DPDP is not something your legal team owns alone: it is a product and data architecture problem. Now.
The Core Tension: Consent and Purpose Limitation vs. AI's Appetite
The DPDP's engine is Section 6: personal data can only be processed with "free, unconditional and clear" consent "for any specific purpose." Not for "AI training." Not for "improving recommendations." For specific purposes. And individuals can withdraw consent anytime.
This creates a hard math problem. A $50M revenue DTC brand we spoke with relies on cross-device user matching to drive repeat purchase models. Under DPDP, that matching is purpose-limited. They need to ask users: "Can we match your email activity to your mobile app activity to predict which category you'll buy next month?" Most won't click yes. The consent rate will crater. Training data will shrink 60-80%. The model's accuracy tank.
Some mid-market operators are hoping for exceptions. The Act does carve out limited cases: "voluntary sharing" and State benefit provision. But for commercial AI models, consent is the gate. The phrase "purposes compatible with the original purpose" does not appear in the DPDP as it does in some European frameworks. The Act is strict.
Data Minimization: The Constraint You Cannot Fudge
Article 5 of the DPDP mandates data minimization: collect "only so much personal data as is necessary" for that purpose. For AI, this is architecturally painful. Large language models and recommendation systems are trained on vast context. The DPDP says no. Collect only what is necessary.
If your recommender needs a user's purchase category, gender, age band, and location to work, you are compliant. If you add search history, email engagement, and device type because "it might help the model," you are not. Data minimization is a hard constraint, not a preference.
The EY survey found 70% of respondents lack familiarity with the Act. That gap shows up in data warehouses. Teams are copying entire customer 360s into AI training sets because it is simpler than defensible data lineage. Under DPDP enforcement, that becomes litigation risk.
Significant Data Fiduciaries and Hidden Compliance Costs
The Act defines "Significant Data Fiduciary" (SDH): an entity whose data processing "has reasonable likelihood of impact on democratic processes, right to freedom of expression, or privacy of individuals." The government has not yet notified the exact threshold (data volume, user count, revenue), but the framework is clear: SDHs face heightened obligations.
Once you are an SDH, you must appoint a dedicated data protection officer, conduct data protection impact assessments, audit your consent flows quarterly, and report material breaches to the Data Protection Board of India. That is not a legal overhead play. That is FTE cost, tooling cost, and operational friction. Most mid-market operators will land in SDH territory if their revenue exceeds $50M and they process personal data at scale.
One enterprise AI team told us their projected DPDP compliance footprint included hiring a compliance engineer (FTE cost ~$80K annually), deploying consent management and data governance tools ($40-60K annually), and running impact assessments ($30-50K annually). For a $100M revenue company, that is real. For a $20M company, it is painful.
Cross-Border Transfers: Your AI Model Is Stuck in India
If you train a model on Indian user data and host it on AWS US-East, you have crossed a border. The DPDP says personal data can only be transferred outside India "under approved conditions." The government retains power to notify restricted countries, and the rules are not yet fully written.
For AI teams, this is an architecture problem. You cannot freely move training datasets, inference pipelines, or model checkpoints across borders without explicit legal gates. Many mid-market operators are operating through India-hosted infrastructure or building consent flows that explicitly exclude non-Indian data processors.
A $40M SaaS company with India-only users faced this directly. They wanted to leverage AWS's latest GPU cluster for model training. Legal said: risky without explicit data transfer governance. They spent two months standing up data localization rules and consent language. Time to value for their recommendation model went from 8 weeks to 16 weeks.
Breach Notification and the Clock You Don't Know
The Act requires personal data breaches be "reported to the Data Protection Board of India and each affected individual in prescribed manner." The phrase "prescribed manner" is where the vagueness lives. The government has not yet clarified the notification timeline (48 hours? 72 hours? 30 days?). Early drafts mentioned 72 hours, but the final rules have not specified.
This matters because your incident response playbook assumes you have time to investigate before notifying. Under DPDP, you may not. Draft your breach protocol now, not after a breach occurs.
What a Mid-Market Operator Should Do Now
DPDP enforcement is phased. The government gave tech giants roughly one year from September 2023 to establish frameworks, so true enforcement pressure accelerates through 2025. But waiting is a cost. Here is what to do now:
First, map your data flows. Which datasets touch Indian personal data? Where does that data live? Who accesses it? Which AI training pipelines ingest it? Do not guess. You will need this map for a data protection impact assessment, and vagueness is liability.
Second, audit your consent language. Pull your current privacy policy and terms of service. Can a user identify the specific purpose for which you are processing their data for AI? If your privacy policy says "improving our products," that is too vague under DPDP. You need granular, purpose-specific consent checkboxes. Assume you will need to roll back old consents and re-collect new ones from users. That is expensive. Do it strategically.
Third, prioritize data minimization in your AI roadmap. When you spec the next model, ask: what data is strictly necessary? Not nice to have. Necessary. Build data lineage tooling so your ML team can justify each feature in training data. Tools like Great Expectations or Collibra help, but they cost money and require process changes.
Fourth, deploy consent management infrastructure. If you are processing Indian user data at scale, you need a system that tracks which consents each user has given, allows withdrawal, and can flag data that should be deleted. Vendors like OneTrust, Transcend, and Snaptik serve this. Cost is $3-10K per month for mid-market scale. Implement it now before a breach surfaces your gaps.
Fifth, engage legal early on architecture decisions. Do not let your engineering team design cross-border AI pipelines in isolation. DPDP cross-border rules will interact with your infrastructure. A 30-minute legal review before you architect saves 12 weeks of rework after a privacy audit.
The Bottom Line: This Is an Execution Problem, Not a Compliance Problem
Many mid-market operators treat privacy law as something you solve in a compliance sprint: update the policy, run an audit, declare victory. DPDP is different. It cascades into your data strategy, your model architecture, your consent flows, your training pipelines, and your cross-border operations. The sooner you treat it as an execution problem, the sooner you will survive it without model paralysis or legal exposure.
The Act's enforcement timeline is still ramping, but the expectation is clear. By late 2025, the Data Protection Board of India will be functional and actively investigating breaches and consent violations. By then, you need to be compliant, not planning compliance.
If you are building AI for mid-market, 10dem works with companies to navigate the data strategy and architecture trade-offs that privacy laws create. The goal is not zero-risk compliance; it is defensible data strategy that lets you build models that perform.

Author
Written by Ankur Garg. Ex-Great Learning and Capital One, with an IIM-Ahmedabad MBA and an IIT-Madras engineering degree. Has built AI products, sold them into enterprises, scaled EdTech from zero, and led P&L, regulatory and BFSI transformation. Advises mid-market and consumer-tech teams on AI strategy, process redesign, and the adoption work that makes AI actually pay off.
Ankur Garg on LinkedIn ↗Want this for your team?
Book a free 30-minute AI opportunity assessment. You'll leave with at least one concrete idea.
Book a call →Discussion
Comments are coming soon.


