Summary
- A developer, known as hashfunction, has shared metadata for approximately 5.6 billion public TikTok videos on Hugging Face, spanning from July 2014 to October 2026, available for free under a non-commercial license.
- The scraping of data violates TikTok's terms, which prohibit automated data collection without explicit permission, and the developer stated the data was obtained from TikTok's private mobile API without requiring a login.
- The dataset directs users interested in commercial use, creator profiles, and daily updates to datasocial.ai, which also offers the scraper's source code for $1,699.
A developer using the handle hashfunction has made available metadata for around 5.6 billion public TikTok videos on the open-source AI platform Hugging Face. This dataset, published under the datasocial account, covers the period from July 2014 through October 2026.
The dataset consists of metadata rather than video content. Each entry includes captions, hashtags, sound IDs, on-screen text, TikTok Shop product IDs, along with metrics for views, likes, comments, shares, saves, and downloads. The total size of the dataset is 460 GB, organized into monthly Parquet files.
Some fields contain TikTok's own designations, such as whether a video is marked as AI-generated or has been excluded from the For You page. The AI designation is frequently absent for older videos, which originate from an archive.
This dataset is openly accessible, allowing anyone to download it. As of now, Hugging Face reports 1,181 downloads. It can also be accessed directly via https://datasocial.ai/.
Why is this significant? AI developers are keen on such data, which TikTok has made difficult to acquire. The combination of captions, hashtags, sound IDs, and engagement metrics for billions of posts provides valuable insights for models that analyze viral trends, sales potential on TikTok Shop, and linguistic patterns in short-form video content.
In contrast, TikTok offers a much more restricted access route to data. Its Research Tools are available only to approved researchers from academic institutions in specific regions such as the United States, EEA, UK, Canada, and Switzerland, as well as select non-profit organizations within the EU, and require an application process.
Automated data scraping is explicitly prohibited by TikTok's rules. Section 3.4 of TikTok's U.S. terms of service states that extracting data using automated software is not allowed unless explicitly authorized by TikTok.
According to DataSocial's report, the developer accessed TikTok's private mobile API using fabricated device identities that mimicked Android devices, reverse-engineered request signatures, and a forged TLS handshake. The report claims that the system gathered 3.23 billion creator profiles, 5.94 billion videos, and 2.8 billion comments within a three-week timeframe, all without requiring a login or account.
The free version serves as a marketplace, licensed under CC BY-NC 4.0, indicating it is free for personal and research use. However, commercial applications, creator profiles, and daily updates must be directed to datasocial.ai, where the source code for the scraper is also available for purchase.
Legal disputes over scraping practices are already unfolding. Reddit filed a lawsuit against Perplexity and three data-scraping companies in October 2025, alleging that they engaged in an "industrial-scale" operation to extract its content for AI training. A federal judge in July largely declined to dismiss the case, according to Law360.
The ongoing claims include DMCA violations against SerpApi for bypassing Google's anti-bot measures. TikTok is not involved in this case, which pertains to scraping through Google search results rather than a private mobile API.
As for the dataset itself, it includes various fields such as captions, utilized music, and additional indicators. While there is no specific column for creator handles, the dataset contains extensive information, including captions and hashtags. Notably, user IDs are accessible through DataSocial.
A similar dataset of 4.5 billion videos, also reported to be collected from TikTok's mobile API, can be found on Hugging Face as well.
The data is free for non-commercial purposes, while the code used to collect it is priced at $1,699.