PLG Metrics Every B2B Product Team Should Track
Activation matters more than signups, and most B2B teams aren't tracking it.

Most B2B product teams building a product-led growth motion are tracking the wrong things. They watch signups and traffic while ignoring activation, the one stage of the funnel that predicts whether a free user ever pays. ProductLed's benchmark survey found that 58% of B2B SaaS companies now run some form of PLG motion, with 91% planning to increase investment in it. Only 34% of those same companies actively track activation. That gap exists because most teams inherited their dashboards from an earlier era of acquisition marketing and never rebuilt them around what a product-led motion actually needs.
A PLG funnel resembles a marketing funnel if you squint, but the logic underneath each stage differs enough that treating them the same produces bad decisions. Attraction and awareness, measured as unique website visitors, is a volume signal: traffic showed up, nothing more. Acquisition, measured as unique signups, confirms someone started the process. Neither says anything about whether a single one of those people found the product useful.
Activation is where that question gets answered. It marks the specific moment a user first experiences the thing the product actually does for them, not a login, not a pageview. Engagement and retention follow as behavioral proof that the value was real and kept being real, through repeated use, team spread, and depth of feature adoption. Conversion, free-to-paid, is where all of that gets monetized. Per ProductLed's benchmarks, the overall average conversion rate across PLG models sits around 9%.
Keeping these stages distinct matters because the diagnosis changes at each one. A low number at acquisition calls for a different fix than a low number at activation, and conflating the two means solving the wrong problem, often at real expense. PLG isn't a universal fit either, and treating it as one is where a lot of teams stumble: it works best where users can find value without talking to a sales rep, which tends to hold for deals under $10,000 in annual contract value. Above roughly $25,000, with multiple stakeholders and a purchasing committee involved, sales-led logic takes over and a different set of metrics matters more. Activation sits at the center of this framework because everything downstream depends on it, and it happens to be the stage most teams skip measuring altogether.
Activation: the hinge metric most teams skip
Only 34% of PLG companies measure activation, according to ProductLed, despite it being the metric most correlated with eventual conversion. That's not a minor blind spot. Most product teams optimize signup flows and traffic sources while ignoring the one number that would tell them whether any of that traffic is finding value at all.
Activation has a precise definition, and it is never account creation or a first login. It's the first successful use of the feature that delivers the product's core value, and the definition has to be specific to the product. For a CRM, that might be first_deal_created. For an analytics tool, first_dashboard_shared. For a design tool, first_file_exported. Benchmarks for B2B SaaS put activation rates in the 25% to 40% range for most companies, with best-in-class products clearing 70% on their core value-delivering feature.
Onboarding design is part of this metric. It's the mechanism that produces it. The 2024 UserPilot Product Benchmarks Report found that 80% of companies with activation rates above 50% use multimedia in their onboarding flow, which suggests the process itself, not just the target, deserves scrutiny. Time-to-Value works as a companion metric here: leading B2B PLG products get users to value in under an hour, and every point of friction between signup and that moment is a place where conversions leak out quietly.
Airtable's experience, as described by Lauryn Isford on Lenny's Newsletter, shows what happens when a team gets specific about this instead of settling for a vague notion of "engagement." Airtable operationalized activation as reaching multi-user collaboration by week four, then rebuilt onboarding around that milestone, including segmenting users by learning style. The result was roughly a 20% lift in activation. Once a team defines its activation event, the job is to instrument it as a named event, make it the primary onboarding target, and track it as a rate against new cohorts, never against total registered users. That denominator dilutes the signal into meaninglessness, and it's the single most common way activation reporting goes wrong.
The activation-to-adoption gap: why first use is not enough
Activation is a moment. Adoption is a pattern, and confusing the two explains a lot of confused post-mortems on stalled PLG motions. A feature that gets activated at a healthy rate but sees little return usage has a value problem: people tried it and decided it wasn't worth coming back to. A feature with low activation in the first place has a discoverability problem instead, since people never found it or never understood what it was for. These need different fixes, and reaching for "we need a better product tour" before diagnosing which problem you actually have wastes a redesign cycle on the wrong target.
Pendo's research on newly launched B2B SaaS features found adoption rates of 20% to 30% within the first 30 days, with high performers reaching 40% to 50% on core workflow features. Timing matters more than the raw number, though. Amplitude's behavioral analytics research found that features achieving repeated usage within the first 7 days show 3.2 times higher 90-day retention compared to features where repeat engagement is delayed. Whatever happens in that first week does outsized predictive work, which is exactly why a 30-day adoption report arrives too late to change the outcome it's measuring.
Breadth matters as much as depth. Accounts using five or more features per month show retention rates of 92% to 96%. Accounts stuck at one or two features retain at 60% to 75%. That's the difference between a healthy account and one already halfway out the door. Activation without follow-through adoption is a leaky bucket, and the two need separate tracking with separate interventions, because fixing one does nothing for the other.
How feature adoption rate is calculated, and why the denominator corrupts most reports
The formula is simple: take the number of users actively using a feature during a given period, divide by the total number of users eligible to use that feature, multiply by 100. The trouble almost never lives in the numerator. It lives in the denominator, and most companies get it wrong in ways that quietly sabotage every decision built on top of it.
Three mistakes show up constantly, and any one of them is enough to make the number useless. Teams include users who don't even have plan access to the feature being measured, which drags the rate down for no real reason. They include inactive accounts, users who haven't logged in during the period at all, as if a dormant account should count against a feature's usefulness. And they mix account-level and person-level counts without distinguishing between them, which makes it impossible to tell whether "adoption" means one champion at each company or the whole team.
Measure against total registered users instead of eligible active users, and the result is a number too low to act on. A number nobody trusts gets ignored, which defeats the purpose of tracking it. Feature adoption rate, calculated correctly, beats a blanket product adoption rate for one reason: it's specific enough to point at a fix, whether that's retraining the onboarding flow for one feature or reconsidering whether the feature solves a real problem at all. Product adoption rate still belongs in the stack, since it answers the broader engagement question, but it's answering a different question than feature-level data does.
None of this works without clean event data underneath it. A practical instrumentation checklist defines and fires events for user_signup, project_created (the first activation milestone for many tools), team_invite_sent (the trigger for viral loops), usage_threshold_80% (a signal for expansion conversations), and feature_adoption_below_median (an early churn flag). Identity resolution, mapping those product events to CRM accounts, is what makes person-level behavior aggregate into account-level decisions. Skip that step and the data stays stuck in a silo no revenue team can use.
Why B2B activation is a team-level problem, not a user-level one
Consumer PLG playbooks break down fast in B2B, and the reason is structural: a single champion can activate fully while the rest of their team never touches the product. Individual-level activation data can look perfectly healthy in that scenario while account-level feature penetration stays quietly low. A team that only watches user-level numbers will miss this entirely, and watching the champion instead of the account is the single most common measurement mistake in B2B PLG.
Real B2B activation means the target team adopts the software, folds it into daily workflow, and starts collaborating inside it. It does not mean one person finished a signup wizard. Three metrics do the heavy lifting here: Time-to-First-Value, Team Activation Rate, and depth of adoption on the core "aha moment" feature. These predict long-term churn and expansion far better than any user-level metric taken alone, because they capture whether the software is becoming infrastructure for a team rather than a tool for one enthusiast.
B2B product adoption tends to move through four distinct phases: Feature Discovery, Depth Adoption, Team-Level Adoption, and Sustained Adoption. Each phase calls for watching a different signal rather than the same dashboard throughout. Accounts where users regularly touch advanced features, integrations and custom dashboards among them, show notably higher retention than accounts that stay on basic functionality. That pattern, multi-user adoption spreading across departments, is also the behavioral signal that precedes enterprise upsell, which is exactly where a PLG motion hands off to sales-assisted expansion.
Companies running a hybrid growth motion that blends bottom-up adoption with sales involvement tend to hit their net revenue retention targets at higher rates than those running pure PLG. Most companies above $10 million in ARR run some hybrid version of the two motions rather than a purebred one, and the data suggests they're right to. Pure PLG past a certain scale is closer to an article of faith than a strategy. Keeping person-level metrics separate from account-level ones, individual activation apart from team activation rate, individual feature use apart from account feature breadth, is what keeps both expansion opportunities and churn risk visible instead of buried under an averaged-out number.
Product-qualified leads: turning behavioral signals into conversion
Product-qualified leads remain badly underused. Only around 24% to 25% of PLG companies use them, per ProductLed's survey, despite a substantial payoff: companies that use PQLs see free-to-paid conversion roughly three times higher than companies that don't. That gap alone should settle the question of whether PQL scoring is worth building. It is.
A PQL is defined by observed product usage signals, not a demographic profile, a job title, or a company size bracket. It's a behavioral threshold: activation completion, unlocking a high-value feature, hitting a usage pattern, or reaching an "aha moment" specific to that product. None of that scoring works if the underlying event data is messy. A PQL model built on unreliable telemetry ends up chasing users who look active on paper but were never close to buying, which wastes sales effort rather than focusing it.
One structural obstacle keeps showing up, and it's an org chart problem more than a data problem: sales most commonly owns free-to-paid conversion, at 23% of companies, while activation, the leading indicator that should inform that conversion motion, sits with product. That split is a large part of why PQL programs stall even at companies that already have the behavioral data to build one. The signals worth watching for PQL designation track closely with the event taxonomy already covered: crossing the usage_threshold_80% mark, firing multi-user collaboration events, and adopting advanced features.
Churn signals live in adoption data, and they arrive early
Retention numbers are lagging indicators. Adoption metrics are the leading ones that explain them, and a renewal lost in the fourth quarter was very often a churn signal sitting in first-quarter adoption data that nobody read at the time.
The behavioral pattern tends to be predictable: declining logins, workflows started but never completed, shrinking engagement with the features that matter most. These show up weeks or months before a cancellation notice arrives, which means there's a real window to act if someone's watching. That window is exactly why static, weekly cohort snapshots aren't enough on their own. Dynamic segmentation, where a user whose outcome events drop below a defined threshold (say, 14 days of reduced activity) gets automatically routed into a re-engagement segment, closes the timing gap that a weekly report leaves wide open. By the time a report surfaces the pattern, the best moment to intervene has often already passed.
Observed behavior change after an outreach message is not proof the message caused that change, and treating it as proof is a common analytical shortcut worth resisting. The more defensible approach measures eligible-cohort adoption rates and genuine repeated use, not opens and clicks, which are engagement theater more than engagement fact. Any behavioral guidance sent to a user should also be grounded in the current state of their account: plan tier, role, permissions. A user who hasn't touched bulk export because their plan doesn't include it is a different problem entirely from a user who has access and simply never found the feature. Retention improves when a team responds to disengagement before it turns into a cancellation, and the adoption metrics covered earlier in this piece are the early-warning system that makes that response possible.
How AI agents are changing what product teams need to measure
Gartner expects 40% of enterprise applications to embed task-specific AI agents by the end of 2026, up from under 5% in 2025. That shift changes what "usage" even means, and most product analytics tooling isn't built to notice.
Agents create a genuine measurement problem, not a cosmetic one. They interact through APIs, firing events and triggering features continuously without ever opening what a session-based analytics system would recognize as a human session. They don't click through a UI, they don't navigate menus, and they don't produce the behavioral patterns that most existing product analytics tooling was built to detect. Any product with meaningful agent traffic in 2026 effectively has a second user class, and existing activation definitions and user segments will misclassify it if nothing changes.
Even the vocabulary is shifting under this pressure. Wes Bush, who popularized the term "product-led growth" (originally coined by Blake Bartlett at OpenView back in 2016), now describes three distinct phases. PLG 1.0 is the user-led version most of this piece has been describing, PLG 2.0 is agentic, and PLG 3.0 is headless. Fast-growing companies including Lovable, Cursor, Gamma, and Perplexity are already operating inside PLG 2.0. Menlo Ventures' 2025 State of AI report found that 27% of all AI application spend flows through PLG motions, a markedly higher share than seen in traditional SaaS, where that figure sits around 7%.
The stakes here are asymmetric in a way that should worry any team building a PLG motion. An agent that fails at a task doesn't complain or file a support ticket. It quietly routes the work to a competitor that succeeds, often without the human overseeing it ever noticing the switch happened. The cost of non-adoption compounds faster in an agentic context than it ever did in a purely human one. Product teams now also have to help users understand when to delegate work to an agent versus when to work directly in the product, since that guidance, done well or done poorly, produces meaningfully different adoption curves. The practical response is to instrument agent sessions as their own category, define activation events specific to agent behavior, and track agent task completion as an adoption signal that sits alongside, but stays distinct from, human usage data.
Putting the framework together: which metrics belong at each stage
Every metric in this framework answers a specific question. Organizing them by that question, rather than by whichever tool happens to report them, is what keeps a measurement stack coherent instead of just crowded.
At the attraction and acquisition stage, unique website visitors and unique signups are volume signals. They answer one question only: is the top of the funnel working. Nothing more should be read into them, and teams that treat signup growth as a proxy for product health are measuring the wrong end of the pipe.
At the activation stage, three metrics carry the weight: activation rate, benchmarked at 25% to 40% for most B2B SaaS with best-in-class products above 70%, Time-to-First-Value, targeted at under an hour for B2B tools, and the product-specific activation event that defines what "core value" actually means for that particular product. These three together answer whether new users are finding the thing the product exists to do.
At the adoption stage, feature adoption rate, calculated correctly against the eligible active user base rather than total registered users, does the answering. It's the metric that separates a feature people tried once from a feature that became part of how a team works, and it's the one most reporting gets wrong through denominator error alone. Layered against the account-level and PQL signals covered earlier, activation rate, feature adoption rate, and Time-to-First-Value form the core of a measurement stack that predicts what happens next, rather than describing, after the fact, what already did.


