TikTok data scraping legal risks are not settled by one deceptively simple question: “Was the data public?” A recent Reddit post from an alleged SaaS founder, who said TikTok demanded a database takedown, revenue, and scraper source code after the founder indexed billions of public data points, shows why builders need a more rigorous framework before turning platform data into a product.

The post should be treated as an unverified first-person account, not a court filing or a confirmed description of TikTok’s claims. Still, the scenario is realistic enough to be instructive. The founder said they collected public profiles, videos, and creator bio information without logging in, offered a SQL-queryable database, sold scraper code, and initially held roughly 30 million email addresses found in public bios. They argued that no authentication wall, paywall, or technical barrier had been bypassed.

That argument may matter. It is not the whole case.

For SaaS founders, marketers, data teams, and AI builders, the central lesson is this: public visibility, technical accessibility, contractual permission, privacy compliance, and commercial reuse are separate questions. Treating them as identical can turn an interesting data product into an expensive legal and operational liability.

The alleged TikTok letter: what the Reddit post actually tells us

The original post in r/SaaS describes a familiar startup pattern: collect data at scale, normalize it, build a simple interface, and monetize the hard work of making a messy platform searchable. In this version, the product reportedly allowed users to query a vast TikTok dataset with SQL, while the operator also sold source code for the scraper.

The alleged recipient framed the situation as straightforward: TikTok content was publicly accessible, collection occurred without a TikTok account, and no paywall, login, hacking, or security circumvention was involved. The founder also compared their work with Google indexing and web-scale AI training, asking why large platforms and AI companies appear able to collect public web material while a small operator is challenged.

Reddit commenters largely did not endorse the legal conclusion. Their practical advice was more restrained:

  • Jurisdiction could change the analysis dramatically.
  • A lawyer should review the actual letter and the underlying facts.
  • The recipient should not make further admissions or detailed replies without counsel.
  • A cease-and-desist demand may be more than a bluff, even if some requested remedies appear aggressive.

That community reaction is more useful than the thread’s rhetoric. Legal exposure depends on where the operator, users, servers, target platform, and affected individuals are located; how data was collected; what happened after notice; whether technical controls existed; and what was done with the data afterward.

Why public data is not automatically permissionless data

A public webpage is visible to a visitor. That does not automatically mean every person can collect it at unlimited scale, republish it in a new database, use it for outreach, sell access, or train commercial systems without consequence.

The distinction is important because a data product typically introduces new conduct beyond ordinary viewing. A visitor may look at one public creator profile. A crawler may request millions of pages, assemble longitudinal records, extract contact details, infer categories, preserve deleted content, and make the results searchable for paying users. Each step can create a different legal or policy question.

Five separate questions founders should ask

Before describing a collection project as “just public data,” separate the issues:

  1. Access: Was the information genuinely available without authentication, and were technical measures bypassed?
  2. Contract: Did site terms apply, and did they prohibit automated collection, commercial use, copying, or derivative databases?
  3. Intellectual property: Are you copying protected videos, images, captions, interface elements, or an original selection and arrangement of information?
  4. Privacy and consumer protection: Does the dataset contain personal information, sensitive inferences, contact details, minors’ data, or information that can be combined to identify people?
  5. Use and distribution: Are you merely conducting internal analysis, or are you publishing, licensing, selling, emailing, enriching, or allowing third parties to query the data?

The answer may be favorable on one dimension and unfavorable on another. For example, a collector may have a credible argument that it accessed a publicly available page without unauthorized entry, while still facing contract, privacy, copyright, trademark, unfair-competition, or state-law claims.

That is why “I never made an account” is not a universal defense. It may affect whether a clickwrap agreement was accepted, but it does not erase every other possible legal theory.

The computer-access question: what hiQ and Van Buren do—and do not—say

The most cited US precedent in public-web scraping discussions is hiQ Labs v. LinkedIn. In its 2022 decision after remand from the Supreme Court, the Ninth Circuit upheld preliminary relief that prevented LinkedIn from using the federal Computer Fraud and Abuse Act, or CFAA, to block hiQ’s access to publicly available LinkedIn profiles. The court’s analysis is frequently read as support for the idea that accessing public web pages is different from entering a gated computer system. (law.justia.com)

The Supreme Court’s 2021 decision in Van Buren v. United States also narrowed one interpretation of the CFAA. The Court held that a person who has authorized access to information does not necessarily “exceed authorized access” merely by using that information for an improper purpose; the statute focuses on obtaining information from areas of a computer system that are off limits to that person. (supremecourt.gov)

Those decisions matter, but founders often overread them.

What those cases can support

In a US case—especially within the Ninth Circuit—there may be a stronger CFAA argument when the crawler accesses information that is truly public and does not defeat authentication, IP blocking, CAPTCHAs, robots enforcement, technical barriers, or other access controls. The federal statute itself is directed at unauthorized access and exceeding authorized access, not at every use a platform dislikes. (uscode.house.gov)

That can be a meaningful distinction for a research project, a search engine, a competitive-intelligence workflow, or an internal product that uses limited public signals.

What those cases do not guarantee

Neither hiQ nor Van Buren creates a nationwide “public data scraping is legal” rule. hiQ involved a preliminary injunction, a specific factual record, a specific platform, and Ninth Circuit law. It did not immunize all scraping, all replication of content, all sales of databases, or all conduct after a cease-and-desist notice.

The legal posture also matters. Winning an argument that the CFAA probably does not apply is not the same thing as defeating a claim for breach of contract, copyright infringement, trespass, state computer-access laws, privacy violations, or interference with a platform’s business. Litigation costs can be existential for a small company even where the final merits are uncertain.

The contrasting Facebook v. Power Ventures litigation is a warning against simplistic rules. The Ninth Circuit found CFAA and California computer-access liability in a situation involving continued access after Facebook’s demand and measures to block the company, along with conduct that used Facebook users’ credentials. (law.justia.com) The factual differences from a no-login public-profile crawler are substantial, but the case illustrates the general principle: technical barriers and post-notice behavior can radically change risk.

TikTok’s terms still matter, even when an account was never created

TikTok’s current US Terms of Service state that they govern access to and use of the platform, including TikTok websites, software, features, technologies, and related services. They also characterize the platform as intended for private, non-commercial use unless TikTok says otherwise. (tiktok.com)

Whether those terms form a binding contract with a particular anonymous visitor is a fact-specific legal issue. Courts look at notice, assent, interface design, the type of agreement, and the user’s conduct. A person who never created an account may have a stronger argument against account-based terms than a logged-in user who expressly clicked “agree.” But a platform does not have to rely only on a contract claim to object to commercial indexing or redistribution.

For a founder, the practical takeaway is not “terms do not matter unless I click a box.” It is: terms are one risk layer among several, and ignoring them can make negotiations, platform access, insurance, fundraising, and customer diligence harder.

TikTok also provides developer and research tools under their own terms. Those routes can be restrictive, slower, and incomplete compared with a custom crawler, but they establish a more defensible permission trail. TikTok’s Research Tools Terms, for example, describe a separate legal agreement for access to designated research tools. (tiktok.com)

For an early-stage company, official access is not merely an engineering convenience. It can be a strategic choice: lower collection flexibility in exchange for a clearer compliance position.

Copyright: facts may be free, but the surrounding content may not be

The Reddit post’s argument centered on public “data points.” That wording matters because raw facts and original expressive works receive different treatment under US copyright law.

US copyright doctrine does not protect facts merely because someone spent effort collecting them. The Copyright Office summarizes the Supreme Court’s Feist decision as rejecting “sweat of the brow” and requiring a minimal degree of creative originality for a compilation to qualify for copyright protection. A white-pages directory, for example, lacked enough originality in its selection and arrangement. (copyright.gov)

That is helpful only up to a point.

A safer distinction: metadata versus expression

A dataset may include relatively factual fields such as:

  • Public handle or profile ID
  • Follower counts at a given time
  • Public posting timestamps
  • Hashtag frequency
  • Basic engagement totals
  • A creator’s disclosed business category

It may also include highly expressive material such as videos, thumbnails, captions, music, comments, profile text, visual assets, and a platform’s interface or curated arrangement. Copying and republishing the latter can create copyright issues that do not disappear because the source profile was public.

Even a mostly factual database can become riskier when it reproduces a platform’s full records, preserves deleted content, reconstructs a browseable substitute for the original service, or exposes media through an interface that lets customers avoid the source platform.

A strong product design principle follows: collect the minimum data needed for the use case, transform it into original analysis where possible, and avoid reproducing expressive content unless you have a clear right or a carefully reviewed legal basis.

The biggest red flag in the post: millions of creator email addresses

The alleged operator said the database once included around 30 million email addresses extracted from public TikTok bios, but that the addresses were later removed from the site and redacted before reaching the frontend. Removing them was a sensible harm-reduction step. It may not, however, resolve all questions about collection, retention, internal access, prior disclosure, backups, downstream buyers, or data-deletion obligations.

Publicly displayed contact information is still personal information in many privacy frameworks. It may also be especially sensitive in practice because mass aggregation changes the context. An email address displayed by a creator who wants brand inquiries is not necessarily an invitation to be placed in a searchable database, resold, profiled, or contacted by anyone for any purpose.

Collection is not the same as outreach—but both matter

In the US, the federal CAN-SPAM Act governs commercial email. The FTC explains that commercial messages must meet requirements including accurate header information, non-deceptive subject lines, a valid physical postal address, a clear opt-out mechanism, and prompt honoring of opt-out requests. (ftc.gov)

CAN-SPAM does not mean that every email sourced from a public profile is automatically lawful to use for any campaign. It means the sending activity has its own compliance rules. State privacy laws, contract terms, deceptive-practices rules, sector-specific regulations, and the recipient’s location can add further constraints.

For marketing teams, the operational rule is straightforward: do not confuse a visible email address with a consent record. At minimum, establish the source of the address, the intended use, suppression logic, retention period, deletion workflow, and proof that your messages comply with applicable commercial-email rules. Before a campaign ever reaches an API, teams should use email address verification and a documented permission-and-suppression process—not just a larger scraped list.

Why Google and AI companies are not a clean comparison

The Reddit author’s comparison to Google and AI companies is emotionally understandable but analytically incomplete. Search engines, foundation-model developers, platform operators, and small data brokers may all collect public web material, but they are not necessarily doing the same thing in the eyes of the law, courts, regulators, or counterparties.

Google’s search function, for instance, generally directs users back to original pages and operates in a legal environment shaped by decades of indexing disputes, publisher controls, caching practices, robots protocols, licensing agreements, and litigation. A commercial database that extracts platform data, creates a standalone query engine, and sells access can look more like a substitute product or data brokerage operation than search indexing.

AI companies face their own ongoing legal disputes about training data, copyright, contracts, and output. The fact that a well-capitalized company is litigating or negotiating through a contested area does not establish that a founder has permission to do the same. Large companies may have licenses, legal teams, different product architecture, insurance, statutory defenses, or simply a greater ability to survive years of litigation.

There is a second-order lesson here for builders: “Others do it” is not a compliance strategy. At best, it is a prompt to investigate the factual and legal differences. At worst, it is evidence that your product was designed with known risk in mind.

Monetization changes the practical risk calculation

A free research dashboard and a revenue-generating data business may not be legally categorically different in every jurisdiction. But monetization changes the incentives, damages theories, negotiation posture, and reputational stakes.

The original poster said the platform requested approximately $45,000 connected to data and scraper-code revenue. Without seeing the letter, no outsider can assess whether that sum was a demand, a settlement proposal, an accounting request, an estimate of alleged damages, or something else. A demand letter is not a judgment, and it does not prove liability.

Still, selling access creates an easy narrative for an aggrieved platform: the operator did not merely observe public content; they allegedly built a commercial business around the platform’s investment, creator ecosystem, data, and infrastructure. Selling the scraper source code can add another concern because it may be presented as enabling widespread reproduction of the contested conduct.

Revenue is not the only commercial signal

A product can look commercial even before it earns money. Warning signs include:

  • Lead generation or sales prospecting based on personal profiles
  • Paid API keys, premium exports, or gated dashboards
  • Licensing data to agencies, brands, recruiters, or AI vendors
  • Ads, affiliate revenue, or paid memberships
  • Public claims that the database replaces platform search or analytics
  • Offering tutorials that materially enable bulk extraction at scale

Founders should assume that commercial positioning will be reviewed alongside the code. Product copy, sales decks, support tickets, pricing pages, investor materials, and social posts can all become evidence of intent and use.

What to do when a platform sends a demand letter

This is not legal advice, and the right response depends on the letter, jurisdiction, and factual record. But the Reddit commenters were right about the first operational principle: do not improvise a legal defense in public or by email.

A founder who receives a demand should take disciplined, reversible steps.

A practical first-response checklist

  1. Preserve the letter and relevant records. Keep the original message, headers, attachments, URLs, timestamps, server logs, code versions, invoices, customer contracts, and access-control records.
  2. Verify authenticity through independent channels. Do not assume an email or Telegram message is genuine merely because it invokes a company name. Independently locate counsel or the company’s legal contact information.
  3. Pause risky distribution where feasible. Consider temporarily disabling public exports, contact-data access, resale, or features most likely to amplify harm while counsel evaluates the situation.
  4. Do not destroy data or alter records casually. A rushed deletion can complicate preservation duties. Ask counsel how to preserve evidence while limiting further processing or disclosure.
  5. Identify every data flow. Map collection, storage, backups, customer exports, vendors, analytics systems, logs, and source-code repositories.
  6. Assess post-notice access immediately. Continuing to crawl after a direct demand or technical blocking can materially worsen the facts, even if the original collection seemed defensible.
  7. Engage a lawyer experienced in internet, privacy, and platform disputes. A general business lawyer may be helpful, but the intersection of computer-access, contract, copyright, and privacy law is specialized.

The most important step is to get the actual facts in front of qualified counsel before sending a detailed response. A founder may be tempted to argue “public means legal,” but broad assertions can box the company into a position that later evidence does not support.

A safer framework for building data products on platform information

The best response to this controversy is not to abandon all web data projects. It is to design them so that the business depends less on indiscriminate copying and more on legitimate, differentiated value.

Build a data-use decision memo before launch

For each source, write a short internal memo answering:

  • What exact fields are collected, and why is each field necessary?
  • Is each field public without authentication?
  • Are any access controls, rate limits, blocks, or technical restrictions present?
  • What do the platform terms and developer policies say?
  • Does the dataset include contact information, minors’ information, sensitive personal data, or inferred traits?
  • Is the product internal analytics, customer-facing search, lead generation, resale, model training, or something else?
  • What are the deletion, correction, opt-out, and suppression mechanisms?
  • Can a user obtain material value without reproducing the original platform’s content?

This document will not magically create legal permission. It does force a team to expose assumptions early, makes outside legal review more efficient, and provides evidence that the company treated risk seriously.

Prefer differentiated outputs over raw mirrors

The least defensible version of a data product is often a full mirror: “Here is another interface for someone else’s content.” More defensible product value can come from analysis that is not a substitute for the source platform, such as aggregated benchmarks, trend detection, creator-selected directories, consented data, first-party surveys, model-generated classifications subject to review, or customer-owned data enrichment.

A healthy test is whether the product still has value if you reduce raw retention. If the answer is no, the company may be monetizing access to the source more than it is creating an original product.

Treat privacy operations as a feature

Teams building creator, influencer, prospecting, or audience products should build removal and correction workflows from day one. Make them easy to locate, keep a log of requests, propagate changes through exports and downstream systems where required, and explain what data is shown and why.

This approach is not only defensive. It makes a data product more trustworthy to enterprise customers, who increasingly ask about provenance, rights, retention, security, and deletion controls during procurement.

The real business lesson: legal ambiguity is a product risk

The r/SaaS thread has the familiar energy of a founder who believes the law is unfairly applied to small companies. There may be real policy questions there. Web scraping law is fragmented; platforms can exercise enormous control over publicly visible information; and courts have not drawn one simple bright line that works for every kind of public data collection.

But founders cannot build durable businesses around a hoped-for double standard. The relevant question is not whether a platform, search engine, or AI lab has done something superficially similar. It is whether your company’s specific collection method, technical behavior, data categories, contracts, use case, and post-notice conduct can withstand scrutiny.

The strongest data businesses plan for that scrutiny before the first customer pays. They use authorized APIs where possible, minimize personal data, avoid publishing raw contact databases, preserve provenance, respect technical boundaries, provide removal processes, and invest in original analytical value rather than building a mirror of someone else’s platform.

In short: public data can be useful input. It is not automatically a free asset, and it is rarely a complete legal answer.

FAQ

Is scraping public TikTok data legal?

It can be lawful in some circumstances, but there is no universal rule. The analysis depends on jurisdiction, whether data was truly public, technical barriers, platform terms, copyrightable content, personal-data handling, commercial use, and conduct after notice. The Ninth Circuit’s hiQ decision is important US precedent, but it is not blanket permission for every scraper or data business. (law.justia.com)

Does not having a TikTok account mean TikTok’s terms do not apply?

Not necessarily. Lack of an account may affect whether someone accepted an account-based agreement, but it does not eliminate other potential claims or make all collection and commercial reuse permissible. Site terms, notice, interface design, and the user’s conduct all matter.

Can I use email addresses displayed in creator bios for marketing?

Visible email addresses are not the same as universal marketing consent. Commercial email in the US must comply with CAN-SPAM requirements, including accurate sender information, an opt-out method, and prompt honoring of opt-outs. Privacy, contract, and state-law issues may also apply. (ftc.gov)

What should I do after receiving a platform cease-and-desist letter?

Preserve the letter and records, verify the sender independently, pause further risky distribution or collection where appropriate, map where the data went, and consult a lawyer experienced in platform, privacy, and internet law. Do not make broad admissions or rely on social-media advice as a legal strategy.

Are search engines allowed to scrape data that startups cannot?

Search engines may have different product behavior, legal precedent, publisher arrangements, technical practices, and resources. Their existence does not automatically authorize a startup to collect, reproduce, sell, or redistribute the same material in a different product model.