Strip away the keynote gloss and a practical question remains: what Salesforce work can agents actually do today, in production, for organizations that aren't Fortune-50 labs? Having built and shipped agents, here's our honest inventory — including the caveats vendors skip.

Customer service: the proven ground
Case deflection is where agentic AI is most mature. Agents resolve order status, returns, account questions, and policy lookups end to end, grounded in your knowledge base — Salesforce reports Reddit reached 46% deflection with response times falling from 8.9 to 1.4 minutes. The caveat: deflection quality tracks knowledge quality. A thin or stale knowledge base produces an agent that confidently underperforms, and the agent gets blamed for the library's sins. How we build service agents →
Sales and revenue: fast follower
Inbound lead qualification, meeting scheduling, and quote drafting validated against CPQ rules all work today. The transactional end — agents assembling orders and invoices — arrived with Agentforce Revenue Management. Caveat: quote-drafting agents amplify catalog quality; a clean product model makes them look brilliant, a messy one makes them dangerous. Catalog first →
Field service: the sleeper hit
Schedule-gap filling after cancellations, pre-work briefs technicians hear en route, plain-language dispatcher actions, and post-work summaries — all shipping, all standing on the optimization engine you already run. Caveat: agents schedule against your work types and durations; if those are fiction, agents automate the fiction faster. The field practice →
Documents and contracts: our signature pattern
Agents that trigger Nintex DocGen — quote generated on deal verification, letter sent on gift receipt, work order filed on job completion — then route for signature and file the result. This is where agentic AI meets a decade of document automation discipline, and it's the pattern competitors without a document practice can't ship safely.
Employee-facing: highest adoption, lowest drama
IT and HR service desks, policy questions, executive-assistant patterns — running in Slack, where nobody has to remember to open anything new. These are often the best first agents politically: internal users are forgiving, data exposure is controllable, and wins are visible.
What we still don't recommend automating
How to pick your first use case (a scoring shortcut)
Score each candidate 1–5 on four axes: volume (does it happen enough to matter?), judgment (reverse-scored — how much human discretion does it truly need?), data readiness (is the ground truth clean and reachable?), and blast radius (reverse-scored — what's the cost of a wrong answer?). Anything scoring 16+ is a pilot candidate; anything scoring under 12 goes on next year's list, whatever the enthusiasm around it. This twenty-minute exercise prevents the most common failure we see: piloting the exciting use case instead of the ready one.
What “production” should mean before you claim it
A demo answers questions; production survives adversaries. Before any agent faces customers, ours pass adversarial testing (hostile inputs, prompt-injection attempts, off-topic pressure), permission verification (the agent can reach exactly what its user context allows, nothing more), and a rollback plan (what happens to in-flight conversations if you pull the agent at 2 pm on a Tuesday). If a partner's definition of done is “the demo worked,” keep interviewing — our seven questions are the follow-up.
High-judgment, low-volume work (complex escalations, sensitive communications), anything where the ground truth data doesn't exist yet, and any regulated step that can't be made deterministic. The honest sequence for everything else: scope one use case, fix the data it touches, pilot in 90 days with measured outcomes, scale from evidence. That's the practice — and if the readiness question comes first, the assessment answers it with evidence.


