NOT a furry tiago.zip

datasets

these datasets may contain personal or sensitive information. by downloading or using them, you agree to comply with all applicable local data protection laws. you may not republish, sell, share or train genai models with the data. we are not responsible for any consequences arising from the use of the data.

twitter.cat

collection of over a billion tweets and 3B profiles actively being scraped from Twitter since 2025, for twitter.cat. scrape ongoing

700+ GB

mastodon.parquet

hundreds of millions of posts and profiles actively being scraped from major mastodon instances since 2026. scrape ongoing

30+ GB

bsky.duckdb

hundreds of millions of posts and profiles scraped from bluesky's firehose and API during 2026. scrape ongoing

90+ GB

twitter-typeahead.db

scraped typeahead twitter data during mid 2025, includes 186k topics and 78M users

40 GB

linktree.db

over 14M linktree profiles, including links, bio, and social media pages

176 GB

manifold.zip

data dumps from manifold markets's internal Supabase API, a prediction markets platform, in the form of JSONL files.

5.0 GB