A model's power is in the data; so is its risk.
Legal compliance of training and test data: copyright and TDM exceptions, KVKK/GDPR legal bases, contractual restrictions and output rights.
The four legal layers of a data set
A training data set touches four regimes at once: copyright (reproduction of protected content and the limits of text-and-data-mining exceptions, rights-holders’ opt-out records), personal data (KVKK/GDPR legal basis, purpose limitation and data-subject rights), contract (source platforms’ terms of use, scraping prohibitions) and trade secrets/competition (data containing another party’s trade secret).

Typical scenarios in corporate practice
Training a model with customer data: do the contract and the privacy notices permit it? Productivity analytics with employee data: the limits of proportionality and transparency. Sets compiled from the web: source-based rights analysis and opt-out screening. Purchased sets: a licence agreement that guarantees the vendor’s chain of rights.
The output side
Copyright protection in generated content depends on human contribution; tool agreements may allocate output rights differently. A regime of filtering, logging and human approval must be set up against outputs that infringe third-party rights — together with the AI Act transparency rules.
From inventory to a defensible chain
The work runs in five steps. Data-set inventory: which sets, from which sources, at what volume. Source classification: each set is tagged against the four regimes — copyright, personal data, contract, trade secret. Legal-basis analysis: KVKK and GDPR bases are determined; on the copyright side, the gap between the text-and-data-mining exceptions of the EU DSM Directive (2019/790) and the absence of an express TDM exception in the Turkish Copyright Act (Law No. 5846) makes a contractual permission architecture necessary for the Türkiye leg. Document production: data cards, supplier warranties, opt-out screening logs. The final step is monitoring: the inventory lives as new sources are added, following the methodology of our Data focus.
Who engages us, and what is delivered?
Typical clients: technology companies training their own models, publishers and platforms looking to license their archives to AI developers, enterprises purchasing ready-made sets, and HR or operations teams building productivity analytics. In German-linked groups, the two regimes (KVKK + GDPR) are managed in a single inventory, and a transfer layer is designed separately for sets shared with headquarters. The deliverables are concrete: a source-based risk map, a legal-basis table, and data-training clauses to be written into procurement and customer contracts. For transparency obligations on the model side, the work is aligned with the AI Act timeline of our Artificial Intelligence focus.
We are by your side for AI Data Set, Copyright & KVKK/GDPR
We produce a data-set inventory and a source-based risk map, align the KVKK/GDPR legal-basis architecture with your compliance programme, and write data-training clauses into procurement and customer contracts. The goal: a defensible data chain, without stopping the model.

Other Applications of This Service
AI Compliance & Governance — our other specialised solutions in this area.
Matter Connections
The focus areas, practice areas, desks and legislation connected with this sub-service.
Our Matters in This Service
The anonymised examples of our work that relate to this service.
The Team Delivering This Service
With our multilingual team of lawyers, well-versed in Turkish and German law, we are by your side.
Related Publications
Fresh perspectives and guides from the Knowledge Centre.
Conditionally: the scope of copyright exceptions, opt-out records, website terms of use, and personal data rules must be analysed on a source-by-source basis. “Everyone does it” is not a legal basis.
Yes, if genuine anonymisation is feasible; however, technical testing is essential and your contracts may provide otherwise. In most cases, a contract update + transparency is the cleanest path.
Depending on how it is used, your company may be held liable; the tool provider’s commitments (IP indemnity) should be sought in the contract. Filters and orderly record-keeping reduce the risk.
AI Data Set, Copyright & KVKK/GDPR — get the right legal support.
Let us identify the right solution together, drawing on our experience in Türkiye and the DACH region.




