A user identified only as hashfunction uploaded a collection of metadata covering about 5.6 billion publicly available TikTok videos to the AI-focused repository Hugging Face. The files span from July 2014 to October 2026 and are stored as monthly Parquet archives totaling roughly 460 GB. Each record lists the video’s caption, hashtags, associated sound identifier, on-screen text, TikTok Shop product and seller IDs, as well as counts for views, likes, comments, shares, saves and downloads. Some rows also include TikTok’s internal flags indicating whether a video is marked as AI-generated or suppressed from the For You feed.
Collection Method
According to a write-up posted on the datasocial.ai site, the scraper accessed TikTok’s private mobile API by emulating Android devices, reverse-engineering request signatures and spoofing TLS handshakes. No user credentials were required. The operation reportedly harvested 3.23 billion creator profiles, 5.94 billion videos and 2.8 billion comments within three weeks. While the dataset on Hugging Face is openly downloadable, the same source notes that the full codebase can be purchased for $1,699.
Legal and Licensing Context
TikTok’s U.S. terms of service expressly forbid automated extraction of data without written permission, a clause cited as Section 3.4 of the agreement. The platform does provide a limited Research Tools program for qualified academic institutions in certain regions, but access is gated by an application process. The newly released dataset is distributed under a Creative Commons Attribution-NonCommercial 4.0 license, allowing free use for research and personal projects. Commercial applications, as well as daily updates and creator profile data, are directed to datasocial.ai, which markets them as a paid service.
Potential Uses and Implications
The breadth of the metadata makes it valuable for developers building models that predict viral trends, optimize TikTok Shop sales, or study short-form language patterns. Because the dataset is free for non-commercial research, it lowers the barrier for academic investigations into content dynamics and recommendation algorithms. However, the method of acquisition raises concerns about privacy, platform policy compliance, and the precedent it sets for large-scale scraping of user-generated content.
Why it matters
The release illustrates a growing tension between open-source AI development and the enforcement of platform terms designed to protect user data and proprietary ecosystems. While the dataset fuels innovation in content-analysis models, it also highlights the legal ambiguities surrounding mass data harvesting and the potential for future litigation, as seen in related cases involving other social media platforms.




