The short version
Traditional marketing automation executes a workflow you drew: if this, then that. Agentic software is handed an outcome instead — reduce week-four churn, fill the trial funnel — and works out the steps itself, calling tools as it goes, reading what happened, and revising. That loop is what the adjective agentic points at.
What the word does not tell you is how much rope the software has. A system can plan five steps and still stop for your approval before it sends anything. That distinction gets lost in most descriptions of the category, and it is the one that decides whether a tool is safe to put in front of your customers.
Nobody neutral has defined it
There is no reference definition of agentic marketing. Wikipedia has no article for the phrase. No standards body has defined it. The definitions that circulate are almost all published by companies selling the software they define — Salesforce, Adobe, Braze and HubSpot each publish one, and so, in this paragraph, do we. That is worth saying plainly rather than pretending otherwise.
The two analyst definitions worth anchoring on are narrower than the vendor ones. Gartner describes agentic AI as systems that autonomously plan and take actions to meet user-defined goals. Forrester's is stricter and more useful, because it ends with the qualifier everyone else drops: software that can flexibly plan and adapt to resolve goals by taking action in its environment, with increasing levels of autonomy. Increasing levels. Not a switch.
The autonomy ladder
Analyst definitions agree the software must plan and act, so the label only reaches the top two rungs of this ladder. What it does not settle is the step that matters most in practice — whether the thing stops for you. These five rungs sort products by how much is decided without you: rules, generative assist, copilot, agent with approval, and unattended.
Rungs one and two decide nothing at run time. Rung three reasons but stops short of acting, which is the line our guide to agents and chatbots draws as well. Only rungs four and five both plan and act, and describing anything below them in the language of the top is what Gartner named agent washing: the rebranding of existing products, such as AI assistants, robotic process automation and chatbots, without substantial agentic capabilities. In the same June 2025 note, Gartner estimated only about 130 of the thousands of vendors claiming the label were genuine, and predicted that more than 40% of agentic AI projects would be cancelled by the end of 2027 — a forecast about 2027, not a measured failure rate.
Rungs four and five are the same software with the checkpoint moved, which makes the step between them an operating decision rather than a purchase — human-in-the-loop marketing is where that choice belongs. The ladder also measures only independence, not blast radius: auto-shipping one in-app tooltip and auto-shipping a broadcast to your whole list sit on the same rung and carry nothing like the same risk. Scope is the second dial, and the one worth being strictest about.
Why the top rung is rare
Long autonomous chains appear to fail in a specific way: the failure rate compounds with every step. Fitting a model to METR’s suite of research-engineering tasks, Toby Ord found the data is explained by an agent having a roughly constant chance of failing per minute of equivalent human work, which produces an exponentially declining success rate as tasks get longer. Each system effectively has a half-life. Ord is explicit that whether this holds on other task suites is still an open question — but the shape matches what practitioners report: short tasks land, long unsupervised ones decay.
Consistency is the harder problem. The tau-bench benchmark, published by Sierra in 2024, measured not just whether an agent completed a task but whether it did so repeatedly: leading models of that generation passed under half the tasks, and succeeded on all eight attempts at the same task less than a quarter of the time in the retail domain. Models have improved since, but the shape of the finding is what matters: even a hypothetical agent that works four times in five is not one you leave alone with your list.
Worse, failures are not always visible. A 2026 study of agent trajectories found that a large share of failures were silent: the agent reported the task complete while the environment showed otherwise, in 45 to 48% of failures in the single-control domains of one benchmark. The same paper found 3% in a dual-control domain, so the rate varies sharply with setup — but silent failure at any of those rates is an argument for an approval step and a decision log, not just for better models.
There is also a security shape specific to marketing. Simon Willison's lethal trifecta describes the danger of combining access to private data, exposure to untrusted content, and the ability to communicate externally — and a marketing platform holds all three by design: your customer list, inbound replies you did not write, and a send button. That combination deserves a checkpoint on its own merits.
How to test a claim
You can place any product on the ladder with a handful of questions. The answers are usually in the documentation rather than the landing page.
| Ask | What a substantive answer sounds like |
|---|---|
| What does it create without me? | Named objects in the platform — a segment, a journey, a test — not a draft in a chat window. |
| Where is the checkpoint, and can I move it? | Approval is configurable per surface, and the default is on. |
| What happens when a step fails? | It stops and says so. If it cannot tell you what failed, it cannot be trusted to run unattended. |
| Can I see why it decided that? | A decision log with the reasoning, not just a timestamp and a status. |
| What does a run cost? | A real number. Multi-step agents call models repeatedly; cost per outcome is a design constraint, not a footnote. |
Now the awkward part, since this page is published by a vendor and the test above is one we wrote — as is the ladder. fromHello is designed for rung four: the specialists create real objects, hold at a checkpoint you can move per surface, and log the reasoning. But none of that is verifiable from outside today. Three of the five answers are design claims and nothing more; the other two — what happens when a step fails, and what a run costs — we cannot make at all, because the hosted product is in early access and pricing is announced at general availability. A test whose model answer happens to describe its author deserves exactly the suspicion this page has been recommending. Give us less credit accordingly, and hold us to the same five questions when it ships.
Two more things are now table stakes rather than nice-to-haves. Disclosure is already law in the EU: the AI Act's Article 50 transparency obligations, which require people to be told when they are interacting with an AI system, have applied since 2 August 2026 and are not limited to high-risk systems. And accountability does not transfer to the software: in Moffatt v. Air Canada, British Columbia's Civil Resolution Tribunal held the airline responsible for wrong information its own chatbot had given a passenger, rejecting the argument that the chatbot was a separate entity answerable for itself. Both of these land on you as the deployer, not on the vendor you are assessing.
What it means for a small team
The category's writing is aimed at enterprises, and the survey data behind it is too: Gartner's 2026 CMO Spend Survey drew on 401 marketing leaders, the vast majority at companies above a billion dollars in revenue. Adoption also runs behind the noise, though the measured figure is about technology leaders rather than marketers: the 2026 Gartner CIO and Technology Executive Survey put AI-agent deployment at 17% of organisations, against more than 60% expecting to deploy within two years.
For a two-person company the practical read is narrower and more encouraging. You do not need the category; you need specific work done — the funnel instrumented, the onboarding sequence written, the churn cohort found. Rung-4 software does that work and asks before it ships. The question of whether it replaces the people who would otherwise do it has its own answer in can AI replace your marketing team, and the shape this takes when the work is split across specialists rather than handled by one general agent is an AI growth team. The underlying software behaviour is autonomous marketing; the difference between an agent and the chatbot it is often confused with is in AI marketing agents vs chatbots.