We build Claude-based agents with real tools, real observability, and real evals. Your team owns the codebase. Your security team can audit it. Your CFO can see the cost dashboard.
Projects are scope-dependent. Free discovery call.
Agent run · live
TriggerUser goal received
PlanDecompose into steps
ToolsCall APIs + search
VerifyCheck result, retry
✓
DoneTask completed
Why this matters
Most AI agents do not survive contact with real users.
The demo runs on a curated input. Production runs on whatever your customers type
at 2am. Tool calls fail silently, prompts drift, costs spiral, and the only person
who understands the prompt has left the company. We build agents that survive that
reality, with the observability, evals, and runbooks to prove it.
What an agent does
Watch one of ours think out loud.
Production agents call tools, surface intermediate results, and answer with citations.
Below is a real turn from a metrics agent we built for a SaaS customer. Every tool call is
observable. Every token is auditable. Scroll up and back down to replay.
No black-box prompt files. No undocumented tool wiring. Every agent we ship comes with the eval suite, the runbook, and the cost dashboard your team needs to operate it after we leave.
01
Tool design before model selection
Most AI agent failures are tool failures. We design the tool surface first, write JSON schemas the model can actually call reliably, and only then pick the model that fits.
→ Tool calls that fail loudly and retry, rather than silently returning nothing.
02
Skills, not megaprompts
Domain knowledge lives in versioned, testable skills the agent loads on demand. Your prompt stays under 200 lines. Updating a workflow does not require redeploying the agent.
→ Iteration cycle drops from days to minutes.
03
Observability from commit one
Every tool call, every token, every retry traced to OpenTelemetry. Cost and latency dashboards live before the agent talks to its first user. Debug is grep, not vibes.
→ Mean time to debug a failed run under 10 minutes.
04
Guardrails the legal team accepts
Input validation, output filtering, prompt injection defenses, scope limits enforced in code. Your security team reviews the agent the same way they review any service.
→ Passes SOC 2 and enterprise procurement review.
05
Prompt caching wired by default
System prompt and skills cached. Cache hit rate measured per conversation rather than assumed. Repeat context is not paid for twice, which is where most of an unoptimised token bill goes.
→ Cost per session measured and budgeted before launch.
06
Handoff documented for your team
Every agent ships with a runbook, a prompt change checklist, an evals suite, and onboarding docs. Your team owns it after week 12, not just operates it.
→ No vendor lock-in, no consultant dependency.
Per session
agent cost logged per session, against a ceiling you set
Prompt caching and on-demand skill loading keep it down, and a per-tenant cap means finance knows the maximum before the month starts rather than after.
The eval layer
Tests that catch prompt regressions before users do.
Every agent ships with a Vitest eval suite mapped to real user cases. CI runs it on every prompt change. Cache utilization is asserted, not assumed. Regressions block merge.
evals/support-agent.eval.tsts
1// evals/support-agent.eval.ts2import { describe, it, expect } from 'vitest';3import { supportAgent } from '../agent';4import cases from './fixtures/support-cases.json';56describe('support agent', () => {7 for (const c of cases) {8 it(c.name, async () => {9 const result = await supportAgent.run(c.input);10 expect(result.tool_calls).toContainEqual(11 expect.objectContaining({ name: c.expected_tool })12 );13 expect(result.message).toMatch(c.expected_pattern);14 expect(result.usage.cache_read_input_tokens).toBeGreaterThan(0);15 });16 }17});
Process
How an agent project runs.
01
Discovery
Two weeks. We map the workflow, identify the tools the agent needs, define the success metric, and lock the eval set. You see a paper prototype before any code.
→ Fixed scope, fixed price, no surprises.
02
Build
Four to six weeks. Tools first, then prompt, then skills, then guardrails. Staging deploy by week three. Eval suite runs on every commit from day one.
→ You can talk to the agent in week three.
03
Launch + monitor
Two weeks. Canary rollout, observability dashboards live, on-call coverage during the first 30 days. Handoff docs and team training before we step back.
→ Your team owns the agent at week 12.
What it costs
Agent pricing follows how much the agent is allowed to do.
A read-only assistant over your docs and an agent that takes actions in production are the same sentence and very different projects. Here is what moves the number.
What moves the number
What it can touchReading is cheap. Every action an agent can take needs permissions, a confirmation path and an audit trail, and that is most of the build.
How many toolsEach tool has to describe itself well enough that the model picks it correctly. The count matters less than how carefully each one is written.
What happens when it is wrongRetries, fallbacks and a human handoff. An agent with no failure path is a demo, not a system.
How much context it needsFeeding it your data means deciding what it sees, how fresh that is, and what it must never see.
Who watches it runTracing every run so a failure names the step that failed is work, and it is what makes the thing maintainable.
Small, well-defined work is usually better handled as an hourly block than scoped as a project, and we will say so rather than inflate it into one.
Rough idea to delivery
You send the detailsWhat you want built, roughly, plus budget range and timing. The form asks for all of it
Within 4 business hoursWe read it and reply. Nothing is scheduled before we know what it is about
Then the callFree discovery, booked once there is enough on the table to make it worth your hour
Within 48 hours of the callA written fixed-price quote you approve before anything starts
Eight to twelve weeks for a production agent with a real tool surface. Discovery and tool design takes two weeks, the agent itself ships in four to six, evals and hardening run in parallel, and a soft-launch monitoring window closes it. Faster if you already have the tool APIs.
Why not just use ChatGPT or a no-code agent platform?
No-code platforms work for demos. They fall over at production traffic, custom tools, observability requirements, and enterprise security review. We build agents your security team can audit and your engineering team can own.
Which models do you use?
We default to Claude Sonnet 4.5 with the Claude Agent SDK because the tool calling and skill system fit production workloads. We also ship on OpenAI and open-weight models when latency, cost, or data residency requires it.
How do you handle prompt injection and abuse?
Input validation against a schema, allowlist for tool inputs, output filtering for sensitive data, scope limits on what the agent can read or write, and a rate-limited, logged audit trail. We include a red-team pass before launch.
Will the agent be expensive to run?
Not if it is built right. Prompt caching, skill loading, and tool design typically keep production cost per session under 5 cents. We give you a cost dashboard and an alert before launch so surprises are impossible.
What does it cost?
Agent work is estimated against scope rather than sold at a list price. A read-only assistant over your documentation and an agent permitted to act in production are very different projects. The section above sets out what moves the number.
WBComDesigns is amazing. The WP plugins and themes are of top notch quality with lots of features and updated very frequently. Supports are very responsive and helpful. I use Reign with TutorLMS, and the experience is surprisingly smooth.…
Hoang Phan·VN·
The job well done!
The job well done, well in time - Thanks so much Anmbiya and team!
Happy Client·ZA·
WBCOM Designs is Amazing!
WBCOM Designs has constantly updated and supported their plugins and themes. Recently the new BuddyNext project is Amazing! - and the new line of addons like Mediaverse, Jetonomy, Mediashield, and Listora.. are likewise Amazing! - and very…
Edward S.·US·
High quality WordPress plugins and themes
We’ve hosted a number of Wbcom Designs clients on Levamo, and our experience with their products has been very positive. Their plugins and themes are well built, their support has been reliable, and they seem to be very innovative and…
Michael Eisenwasser·US·
I have worked with Wbcom Designs on…
I have worked with Wbcom Designs on several custom development requests for my WordPress platform, and my experience has been excellent.
christian nicolas·BJ·
Love this company
Love this company. Loved the REIGN product. They were there the entire time helping me set things up; and when things went wrong? They were quick to take action and help find a solution.
Fable Fortitude·CA·
Seriously, one of the best "software tech experiences" I've ever had!
After 16 years of buying WordPress themes and plugins, I know exactly what bad support looks like and Wbcom Designs is the polar opposite. My setup was a nightmare: multiple tools, deep integrations, custom configurations that required…
Duston McGroarty·US·
I was using an excellent plugin created…
I was using an excellent plugin created by Wbcom Designs and had both an error and discovered a slight bug in one aspect of the plugin. After creating a support ticket - I got a super-quick response and discovered the error was on my part…
Edward Bonthrone·US·
Excellent Theme, Powerful Plugins and Outstanding Support
I am using the REIGN theme and several plugins from Wbcom Designs on my website. The theme is beautifully designed, and the plugins are user-friendly. Everything works smoothly, and the features are perfect for building professional…