← Blog

The Hidden Cost of "Free" AI Notetakers: Where Your Meeting Data Actually Goes

    A free AI notetaker does not cost nothing. It costs the recording, transcript, and analysis of your most sensitive business conversations, handed to a company whose economics you have not examined. That is not an accusation of wrongdoing. It is an observation about how the deal is structured. When a product is free and it captures deals, pricing, personnel discussions, and customer confidences all day, the sensible question is not "is this too good to be true" but "what do the terms I agreed to actually allow, and where does this data physically and legally end up?"

    This post answers that using the vendors' own privacy policies, with sources so you can verify every claim yourself rather than take my word for it. Policies change often, so the point is less any single fact and more the habit of reading them. Everything below reflects the policies as published in mid-2026; check the effective date on whatever you use.

    The frame

    With a free notetaker, the meaningful cost is not money. It is three things the terms of service govern: whether your conversations can be used to train models, where your data is processed and under whose law, and how long the record is kept. Read for those three, not the price.

    "Free" means the business model is somewhere else

    Software that records, transcribes, and analyzes meetings is expensive to run: storage, compute, and increasingly large-language-model inference on every call. A vendor giving that away at no charge is funding it somehow, whether through a paid tier you are meant to upgrade to, through the strategic value of the data itself, or both. None of that is inherently sinister. But it means the free tier is rarely where the vendor's incentives point, and the free tier is usually where the data protections are thinnest. The specifics live in the policy, and the policies differ more than the near-identical marketing pages suggest.

    Whether your meetings train the AI

    This is the claim that varies most between vendors, and where reading the actual policy matters most. There is a critical distinction to hold: a vendor training its own models on de-identified data is a different thing from a vendor letting a third-party model provider train on your content. Most now contractually prohibit the second. The first is where they diverge.

    Otter.ai is the clearest example of a policy that permits training on your content. Its privacy policy, effective June 16, 2026, lists among its purposes "training our proprietary AI technology on de-identified audio recordings and on transcriptions (which may contain Personal Information)." Read that carefully: it says Otter trains its own AI on de-identified audio and on transcriptions, and it acknowledges those transcriptions may still contain personal information. Whether an opt-out exists specifically for training, and whether it is on by default, is not something I could confirm from the policy text itself, so I will not assert it. Verify it against your own account settings and the current policy at otter.ai/privacy-policy before relying on either answer.

    Fathom takes a middle path its policy states plainly: it uses de-identified data "to improve our Services by training, improving, and customizing our in-house artificial intelligence models," configurable in account settings, while prohibiting third parties such as OpenAI, Anthropic, or Google from training on your content. So: own models yes, third-party models no, with a stated opt-out.

    On the other side, Fireflies.ai's policy states "we do not use personal information for AI model training and we contractually prohibit our vendors from using this information for their own model training." tl;dv similarly markets a no-training posture, though the explicit commitment sits on its security page rather than in the core policy, so I would cite it as a security-page claim, not a binding policy term. Read.ai ties training to your account settings and a named opt program, without a policy sentence that unambiguously fixes the default. The lesson is not that any one of these vendors is bad. It is that "does this tool train on my calls" has genuinely different answers, and only the current policy tells you which.

    Where the data actually lives, and whose law reaches it

    Most of the widely used notetakers are US companies that, by their own policies, process and store data in the United States. Otter states it relies on cloud providers "including Amazon Web Services, based in the United States." Fathom states its services "are hosted in the United States." Read.ai states it is "based in the United States" with servers "in the United States and other countries." Fireflies is US-based by default and offers regional storage on a paid basis. The clearest EU-default exception among the common tools is tl;dv, whose policy describes primary processing and storage "in facilities located in the European Economic Area," naming German and Finnish infrastructure providers.

    Why does the country matter more than the marketing? Because location of storage and jurisdiction over the provider are two different facts, and the second is the one that decides who can compel your data. A US company is subject to US law, including the CLOUD Act, which lets US authorities require a US provider to produce data it controls regardless of where the servers physically sit. An "EU data center" operated by a US company does not remove that reach. For a European team recording board discussions or a customer's confidential roadmap, that is a specific exposure, not a hypothetical one. We unpack the full residency-versus-sovereignty gap in what sovereign AI actually means, and the same jurisdiction problem is why Gong's EU data handling gets scrutinized by compliance teams.

    There is a second-order point worth naming. Even providers that keep primary data in the EU sometimes route the AI step elsewhere. A policy can promise EU storage while the model that generates your summary runs through US infrastructure or a US-controlled model provider. The processing that matters for surveillance risk, the moment your words pass through a model, can leave the jurisdiction even when the file does not. When you read a policy, follow the AI step specifically, not just the storage line.

    How long the record survives

    Retention is the quietest of the three costs and often the least specific. Some policies publish a clear window. tl;dv's policy sets a free-tier retention for recordings measured in months, with paid accounts kept until deletion, a concrete free-versus-paid difference. Fireflies states it deletes account-related personal information within 30 days of account closure. Read.ai states it stores audio and video "in no case for longer than 2 years." Others, including Otter, keep data "as long as necessary" without publishing a fixed number, which in practice is an open-ended retention you cannot easily bound. Whenever a policy uses "as long as necessary" with no figure, read that as the vendor reserving discretion, and decide whether you are comfortable granting it over recordings of your candid internal conversations.

    These questions are being tested in court

    That this is a live legal area, not a manufactured worry, is visible in public court records. As of mid-2026, several proposed class actions against Otter.ai were consolidated in the Northern District of California, alleging the tool recorded and used meeting data without all-party consent under US wiretap and privacy statutes. Those allegations are unproven, Otter has moved to dismiss, and no ruling had issued at the time of writing. I mention it not to imply liability, which no court has found, but because it shows regulators, plaintiffs, and courts are now actively probing exactly the questions this post is about: consent, training use, and control of meeting data. The prudent response is not alarm. It is to read the terms before you record.

    A short checklist for reading any notetaker's policy

    You do not need to be a lawyer to protect yourself here. Open the provider's current privacy policy and look for four things:

    1. Training. Does it permit using your content, even de-identified, to train the vendor's own models? Is there an opt-out, and is it on by default? Does it separately prohibit third-party model training?
    2. Jurisdiction. Where is data processed and stored, and is the company subject to US law such as the CLOUD Act? Follow the AI step specifically, not just the storage claim.
    3. Retention. Is there a specific window, or only "as long as necessary"? What does deletion actually remove, and how fast?
    4. Sub-processors. Is there a named, current list? Which third parties, and in which countries, touch your recordings?

    If a reassuring claim appears only on a marketing page and not in the binding policy, treat it as marketing. And check the effective date, because the policy that mattered when you signed up may not be the one in force today.

    The honest alternative

    The point of this post is not that you should stop using meeting assistants. The value of good meeting intelligence, durable memory and real coaching, is exactly why we build one. The point is that a record of your most candid business conversations is among the most sensitive data your company holds, and "free" should never be the reason you stop asking where it goes.

    That is the gap Numi is built to close. It is a sovereign meeting assistant that keeps call audio, transcription, storage, and AI analysis under EU jurisdiction on infrastructure we can point to, so the answers to the four questions above are ones we can give plainly rather than bury. If you are comparing options against the criteria in this post, our GDPR-compliant meeting assistant comparison lays them out side by side. Free is a price. It is not a data protection strategy.

    Frequently asked questions

    Do free AI notetakers use my meetings to train their AI?

    It depends on the vendor, and the only reliable source is each provider's own privacy policy. Some policies expressly permit training on de-identified audio and transcripts of your meetings, sometimes on by default with an opt-out you have to find. Others state they do not train on customer content and contractually prohibit their vendors from doing so. The important distinction is between a vendor training its own models on de-identified data and a vendor allowing a third-party model provider to train on your content, which most now prohibit. Always read the current policy rather than the marketing page.

    Where is my data processed when I use a free notetaker?

    Most of the widely used notetakers are US companies that process and store data in the United States by default, per their own policies. EU data residency, where offered at all, is often gated behind a paid or enterprise tier. A few providers default to EU or EEA storage. Because processing location and, more importantly, the legal jurisdiction of the provider determine who can compel access to your data, this is worth checking before you record a single call.

    How long do free notetakers keep my recordings?

    Retention varies widely and some policies do not publish a fixed window at all. Published practices among common tools range from deleting free-tier recordings after a few months, to deleting data within 30 days of an account deletion request, to keeping audio and video for up to two years. If a policy only says data is kept as long as necessary without a number, treat that as an open-ended retention you cannot easily bound.

    How do I actually check what a notetaker does with my data?

    Read the provider's current privacy policy and look for four things: whether it permits training on your content and whether any opt-out is on by default; where data is processed and stored, and whether the company is subject to US law such as the CLOUD Act; how long recordings are retained and what deletion actually removes; and the named list of sub-processors. Policies change, so check the effective date. If a claim only appears on a marketing page and not in the binding policy, treat it as marketing, not a commitment.

    A record of your most candid business conversations deserves a plain answer to where it goes. Numi is a sovereign meeting assistant that keeps audio, transcripts, and analysis under EU jurisdiction on infrastructure you can point to.

    Get Early Access