tiago.zip
these datasets may contain personal or sensitive information. by downloading or using them, you agree to the dataset license. access can be revoked at any time, and all data is provided as-is. we are not responsible for any consequences arising from its use.
collection of over 3 billion tweets and 3B profiles actively being scraped from Twitter since 2025, for twitter.cat. scrape ongoing
3+ TB
dozens of millions of posts and profiles from across mastodon and the fediverse, deduplicated across instances and exported monthly.
50M+ posts
hundreds of millions of posts, profiles and follows scraped from bluesky's firehose since 2026. scrape ongoing
200M+ posts
billions of videos, reposts, and billions of comments scraped from tiktok since early 2026. scrape ongoing
2B+ videos
collection of over 15M gifs and stickers from tenor in video and database form, harvested after the 2026 api shutdown. scrape ongoing
8+ TB
scrape of the cutting room floor's mediawiki after they started banning datacenter ips and cracking down on bots. harvest started september 2026.
70 GB
hundreds of millions of truth social posts from about a million accounts, scraped since early 2026. scrape ongoing
100M+ posts
scraped typeahead twitter data during mid 2025, includes 186k topics and 78M users
40 GB