NOT a furry tiago.zip

datasets

these datasets may contain personal or sensitive information. by downloading or using them, you agree to the dataset license. access can be revoked at any time, and all data is provided as-is. we are not responsible for any consequences arising from its use.

read the full license

twitter.cat

collection of over 3 billion tweets and 3B profiles actively being scraped from Twitter since 2025, for twitter.cat. scrape ongoing

3+ TB

fedi.parquet

dozens of millions of posts and profiles from across mastodon and the fediverse, deduplicated across instances and exported monthly.

50M+ posts

bsky.parquet

hundreds of millions of posts, profiles and follows scraped from bluesky's firehose since 2026. scrape ongoing

200M+ posts

tiktok.parquet

billions of videos, reposts, and billions of comments scraped from tiktok since early 2026. scrape ongoing

2B+ videos

tenor archive

collection of over 15M gifs and stickers from tenor in video and database form, harvested after the 2026 api shutdown. scrape ongoing

8+ TB

the cutting room floor

scrape of the cutting room floor's mediawiki after they started banning datacenter ips and cracking down on bots. harvest started september 2026.

70 GB

truthsocial.parquet

hundreds of millions of truth social posts from about a million accounts, scraped since early 2026. scrape ongoing

100M+ posts

twitter-typeahead.db

scraped typeahead twitter data during mid 2025, includes 186k topics and 78M users

40 GB