NOT a furry tiago.zip

data processing notice

if you'd like to opt out, please check tiago.zip/trawler. you do not need to read the rest of this page first.

This notice explains how the datasets published at tiago.zip/datasets are collected, processed, stored, and made available, and what you can do about material that relates to you. Together those collections are called the Corpus.

This notice covers only the Corpus. It does not cover cookies, analytics, server logs, or anything about how you use the tiago.zip website itself, which are handled separately.

1.Who we are and where we operate

The Corpus is operated by the operator of tiago.zip and twitter.cat, a natural person resident in Portugal, who is the controller of the Corpus where applicable law treats us as such. Where it is necessary to bring or defend legal proceedings, or where a court or supervisory authority requires it, we will identify ourselves in full.

The services run on distributed infrastructure supported by providers in more than one country, so the activities described here may take place in, or involve systems located in, Portugal, Germany, and the United States. We aim to process information consistently with the General Data Protection Regulation, Regulation (EU) 2016/679 (the GDPR), as generally the most protective regime that might apply, without giving up any argument available under any other law.

Questions and requests go to legal@tiago.zip.

2.What this notice covers

We collect, index, aggregate, and process information that was publicly accessible at the time we collected it. This is information people chose to publish, and that they or the platform made visible to the general public without any access restriction. The defining feature of everything in the Corpus is that it was already public when we encountered it. We do not reach behind locks, logins, privacy settings, or any other barrier, and the Corpus is not built from anything that was private at the time of collection.

The Corpus can include:

We do not seek to collect special categories of data under Article 9 of the GDPR. We apply measures intended to avoid targeting sources and fields whose purpose is to reveal such data. Where we incidentally and residually process special-category data despite those measures, it is processed only because the person manifestly made it public within the meaning of Article 9(2)(e). We do not knowingly build the Corpus around sensitive characteristics.

3.Where the information comes from

The Corpus is built from public information published on and made publicly accessible through third-party platforms, currently X (formerly Twitter), Mastodon and other fediverse instances, Bluesky, Linktree, and Manifold Markets. When someone posts publicly, fills in a public profile, or follows an account in a way the platform displays openly, they disclose that information to the public at large, and that already-public layer is what we work with. We did not create this information and are not the original place it was published. We gather, organise, and index what was already made public.

We are not affiliated with, endorsed by, or sponsored by any of those platforms or their operators, and we have no formal relationship with any of them. Trademarks, service marks, and trade names mentioned belong to their respective owners and are used only to describe the source of the information accurately.

4.How we collect it

We gather public information through automated means, on a continuous and recurring basis rather than as a one-time snapshot. Because the Corpus functions as a lasting, searchable record, information we have collected is retained, re-indexed, enriched, and re-processed as part of ongoing operation, and we treat it as part of the Corpus independently of whether it remains visible on the source platform afterward. Later changes at the source do not automatically change our copy, though section 10 explains how removal affects what we hold over time.

5.Why we process it

We process the Corpus to operate a searchable public record of already-public information: to index and organise the material so it can be searched and retrieved, to build aggregate statistics, visualisations, and analytical insights from it, to generate the derived signals described above, and to make it available for research under licence. We also process it to develop, improve, secure, and reliably operate the services, and for further purposes compatible with these.

6.Our legal basis under the GDPR

To the extent the Corpus contains personal data within the meaning of Article 4(1), and to the extent the GDPR applies to a given activity, we rely on our legitimate interests under Article 6(1)(f). Those interests include operating a public search and discovery service, organising and indexing information the people it relates to already chose to make public, enabling public access to information of public interest, and carrying out research and analysis on that information, and enabling others to do the same.

We have carried out and documented the three-part assessment that Article 6(1)(f) requires, covering the purpose, the necessity, and the balancing of our interests against the interests, rights, and freedoms of the people the data relates to, and we keep that assessment under review. Several things matter to our conclusion: the information was deliberately published and made public by the person it relates to; because it was already public, the reasonable expectation of privacy attaching to it is correspondingly limited; we process the minimum necessary to operate the Corpus; and the rights and routes in this notice provide meaningful safeguards. None of this removes your ability to object to processing based on legitimate interests, and section 10 explains how.

Where the Corpus contains special-category data under Article 9, we rely on Article 9(2)(e): such data is processed only where the person manifestly made it public, and only as described in section 2.

7.How the Corpus is made available

Parts of the Corpus are made available to other people, in some cases openly and in some cases only after a request is reviewed and access is granted individually. Everyone who receives any part of the Corpus does so under the tiago.zip Dataset License, which they must accept first.

That licence obliges every recipient, among other things, to:

Recipients act as independent controllers and are responsible for their own compliance. We keep a record of who accepted the licence and when. We rely on infrastructure providers and processors whose systems are necessary to run the services, and we may disclose information where the law requires it.

8.Information crossing borders

Running the services involves systems and providers in Portugal, Germany, and the United States. Personal data within the Corpus may be transferred to, stored in, or accessed from countries outside the European Economic Area. Where that happens and the GDPR applies, we put in place appropriate safeguards consistent with Chapter V, which depending on the recipient may mean relying on an adequacy decision, standard contractual clauses, or another lawful transfer mechanism.

9.How long we keep it

Because the Corpus acts as a historical and analytical record of public information, we keep it for an indefinite period, for as long as it serves the purposes set out here, and subject to the rights in section 10 and to applicable law.

10.Your rights and how to use them

Where the GDPR or a comparable law applies, and subject to the conditions and exceptions in that law, you may have rights over personal data relating to you that sits within the Corpus. These can include access, rectification of inaccurate data, erasure (itself subject to the exceptions in Article 17(3)), restriction of processing, objection to processing based on legitimate interests under Article 21, and, where applicable, data portability.

The fastest route is the opt-out at tiago.zip/trawler. You can also write to legal@tiago.zip. So we can find the right information and confirm the request comes from the person it concerns, include enough detail to identify the material, at minimum the relevant handle or numeric identifier and a description of what is at issue. We may ask for additional information reasonably needed to verify your identity before we act.

We assess and respond to verified requests within the time the law allows, normally one month under the GDPR, extendable by a further two months where a request is complex or numerous. When we remove something, everyone holding a licensed copy is required to delete the matching records too, within seventy-two hours of our notice, as described in section 7. We cannot guarantee that a recipient who breaches their licence will comply, but that obligation is contractual and we enforce it.

In some cases the law lets us decline a request in whole or in part, including where processing is necessary for the exercise of the right to freedom of expression and information, for archiving in the public interest, for scientific or historical research, or for establishing, exercising, or defending legal claims. Where one of those grounds applies we may rely on it, and we will tell you so.

If you are in the European Economic Area and think our processing breaks data protection law, you have the right to complain to a supervisory authority, in particular where you live, where you work, or where you believe the problem occurred. In Portugal this is the Comissão Nacional de Proteção de Dados (CNPD).

11.Children

The services are not aimed at children, and the Corpus is not built by going looking for children's information. We process only material already published publicly, and use of the source platforms is itself subject to their minimum age rules. We do not knowingly seek out or target information from anyone below the minimum age allowed to use those platforms where they live. If we become aware that information relating to someone below that age has entered the Corpus other than through that person's own lawful public use of a platform, we will take reasonable steps to deal with it once a verified request reaches us through section 10.

12.Security

We use reasonable technical and organisational measures to protect the Corpus against unauthorised or unlawful processing and against accidental loss, destruction, or damage, appropriate to what the Corpus is: information that was already public when gathered. Access to the non-public parts is restricted, and distribution is gated behind the licence described in section 7.

13.Changes to this notice

We may update this notice from time to time. The version that applies at any moment is the one marked by the date at the top. We will not notify every individual of every change, so check the date at the top.

14.Contact

Questions about this notice, or requests under section 10, go to legal@tiago.zip. Terms for using the datasets are in the dataset licence.