SSS · AI Data Sets, Copyright & KVKK/GDPR

Can we train a model with data we collect from the internet?

Conditionally: the scope of copyright exceptions, opt-out records, website terms of use, and personal data rules must be analysed on a source-by-source basis. “Everyone does it” is not a le…

Updated · July 20261 min readCategory · AI Data Sets, Copyright & KVKK/GDPR
Short answer

Conditionally: the scope of copyright exceptions, opt-out records, website terms of use, and personal data rules must be analysed on a source-by-source basis. “Everyone does it” is not a legal basis.

Conditionally: the scope of copyright exceptions, opt-out records, website terms of use, and personal data rules must be analysed on a source-by-source basis. “Everyone does it” is not a legal basis.

For EU-facing training there is a specific gate: the text-and-data-mining exception in Article 4 of the DSM Directive applies only where the rightsholder has not reserved use in a machine-readable way, so honouring opt-out signals is part of having a lawful basis at all. And where a source contains personal data, you need a KVKK or GDPR basis for it as well: clearing copyright does not clear data protection, which is a separate gate. That is why the analysis is genuinely source-by-source, and why a blanket scrape is not defensible.

Shall we apply this matter to your situation?

Tell us your specific situation in a few sentences; we'll assess it with the right team.

Get in touch →
This content is for general information only and does not constitute legal advice. Please contact our team for an assessment of your specific circumstances.
Categories
AI Data Sets, Copyright & KVKK/GDPR

The right start means a predictable process.

From the first meeting to completion of the work; let's plan every step transparently.