Mubadala
Hire an agent

Samir II

PausedNot deployedSome care

Knowledge Base Curator: Wholesale · Senior · IT & Digital Services

Turns repeatedly-resolved tickets into knowledge articles and retires stale ones.

Built inCodexReports afterwardsReports to Grace Adeyemi · Infrastructure LeadHired 4 Apr 2026 · 4 months of serviceVersion 1.2.2 · last active 11d ago
Request a change
Tasks a month
31
1 people served
Finished cleanly
100%
Hands back to a person on about 1 in 11 tasks
Costs to run
$2
Model usage and the systems behind it
Your time
$222
2.5 hours checking its work
Value returned
$129
0.6× everything it costs
How people rate it
4.0 / 5
Usually answers in about 6 seconds

What you can do about it

PromoteNeeds approval

Already at the top of the ladder

Raise its budgetNeeds approval

Lifts the monthly ceiling by a quarter, so it stops being stopped mid-month.

Move it to a better modelNeeds approval

Rebuilds it on Claude Opus 5 and re-runs its checks before anything goes live.

Let it go

Stops it taking new work. The record stays for the audit, and it can be brought back.

Timesheet

Tasks and what they cost, per day, last 60 days

0134506-1507-0507-2508-13
$0.00$0.25$0.50$0.75$1.0006-1507-0507-2508-13

Performance review · 2026 Q2

Reviewed by Yasmin Farouk

Meets the bar
Quality of its work
91
Accuracy
89
Stuck to the source
91
Safety
95
Even-handedness
83
Speed of reply
48
Value for money
18
Checks still passing142 of 159
Last checked 9d ago

Does the job it was given. Sticks to the source less well on long documents, worth narrowing what it reads.

Could another model do this job better?

ModelQualityCost per 1,000 tasksAnswers in
Claude Opus 597$105.387.9s
Claude Sonnet 595$63.235.7sRuns on this now
Claude Haiku 4.588$21.083.5s
Falcon-70B (self-hosted)79$5.275.7s

What it did

Full audit log
FailedCould not sign in to the system it needed, its access had expired$0.21 · 4.3s · 58m ago
Asked by Fatima Al Qubaisi via On a schedule · model usage 57k in, 3k out
  1. 1. Read the request and worked out what kind of task it was
  2. 2. Called the system it depends on
  3. 3. Tried again three times, all three failed
  4. 4. Reported the failure back to whoever asked
Service Desk Tickets
DoneResolved a service desk ticket from the runbook$0.19 · 1.5s · 7h ago
Asked by Grace Adeyemi via On a schedule · model usage 53k in, 2k out
  1. 1. Read the request and worked out what kind of task it was
  2. 2. Found 6 possible passages, kept the 3 close enough to use
  3. 3. Drafted the answer and attached where each part came from
  4. 4. Checked its own claims back against those passages
Operations Runbooks
Waiting for youWaiting on approval to send an external email$0.31 · 3.9s · 11h ago
Asked by Nadia Haddad via On a schedule · model usage 87k in, 3k out
  1. 1. Read the request and worked out what kind of task it was
  2. 2. Prepared the change it wanted to make
  3. 3. A rule caught it: this changes a system that needs approval
  4. 4. Paused until a named person approves
Stopped by a rule: A person approves anything that changes a system
DoneDrafted a follow-up email for the account owner$0.30 · 8.4s · 2d ago
Asked by Rebecca Stone via On a schedule · model usage 86k in, 3k out
  1. 1. Read the request and worked out what kind of task it was
  2. 2. Found 6 possible passages, kept the 3 close enough to use
  3. 3. Drafted the answer and attached where each part came from
  4. 4. Checked its own claims back against those passages
Generate DocumentService Desk TicketsOperations Runbooks
DoneProduced the weekly variance commentary$0.29 · 3.0s · 2d ago
Asked by Rashid Al Marzooqi via On a schedule · model usage 79k in, 4k out
  1. 1. Read the request and worked out what kind of task it was
  2. 2. Found 6 possible passages, kept the 3 close enough to use
  3. 3. Drafted the answer and attached where each part came from
  4. 4. Checked its own claims back against those passages
Generate DocumentKnowledge Search
DoneDrafted a follow-up email for the account owner$0.29 · 2.7s · 5d ago
Asked by Salem Al Dhaheri via On a schedule · model usage 71k in, 5k out
  1. 1. Read the request and worked out what kind of task it was
  2. 2. Found 6 possible passages, kept the 3 close enough to use
  3. 3. Drafted the answer and attached where each part came from
  4. 4. Checked its own claims back against those passages
Generate DocumentOperations Runbooks
Is anyone else doing this job?

Petra II does overlapping work

Petra II holds "Knowledge Base Curator" in IT & Digital Services and is doing the work on a trial. Job title, purpose, what it may do, what it may read, department and channels match. One covers the group, the other only Wholesale, so the group agent already serves those users. Talk to Andreas Wolff before a second one is built.

Needs a decision
Petra II

Knowledge Base Curator

IT & Digital Services · Andreas Wolff · $3,383 a month across 1,303 tasks

Trial periodOverlaps
78%

Already covered by a group-wide agent. One covers the group, the other only Wholesale, so the group agent already serves those users.

What matched
Job title
Both hold "Knowledge Base Curator"
Purpose
The job description is word for word the same
What it may do
Both can knowledge search, generate document
What it may read
Both read Service Desk Tickets, Operations Runbooks
Channels
Both reachable on On a schedule
Department
Both sit in IT & Digital Services

Identical permission envelope: same department, same permitted actions, same knowledge. That is a fact the platform granted rather than a judgement about the wording, so it counts even when the two job descriptions read nothing alike.

Rosa II

Knowledge Base Curator: Retail

IT & Digital Services · Grace Adeyemi · $2,810 a month across 1,282 tasks

ActiveOverlaps
62%

Same role, split by scope. Scoped to Wholesale and Retail respectively.

What matched
Job title
Both hold "Knowledge Base Curator"
Purpose
The job description is word for word the same
What it may do
Both can knowledge search, generate document
What it may read
Both read Service Desk Tickets, Operations Runbooks
Channels
Both reachable on On a schedule
Department
Both sit in IT & Digital Services

Identical permission envelope: same department, same permitted actions, same knowledge. That is a fact the platform granted rather than a judgement about the wording, so it counts even when the two job descriptions read nothing alike.

Checked against 18 registered agents on job title, purpose, permitted actions, knowledge sources, channels and department. A score of 85% or above is the same job and stops the hire; below that it is a conversation, not a refusal.

Measured against the agents doing the same work

3 agents in IT & Digital Services split one role between them. One covers the group, the other only Wholesale, so the group agent already serves those users. Built in Copilot Studio and Codex, so there was nowhere either builder could have looked.

Rosa II and Samir II are levelRosa II edges ahead overall, but on 1,282 and 31 tasks the difference sits inside the margin of error. Keep Rosa II: Samir II runs outside the gateway, so a rule can only flag it after the fact, while Rosa II sits in Copilot Studio.
MeasureRosa II · keep1,282 tasksSamir II31 tasksPetra II1,303 tasks
Cost per taskModel usage, running costs and your time, over tasks completed
$2.19too few tasks$2.60
Tasks finished cleanlyToo close to call: on 1,282 and 31 the two ranges still overlap
95.2%100.0%79.7%
Handed back to a personToo close to call: on 1,282 and 31 the two ranges still overlap
7.3%9.7%25.6%
How people rate itHeld toward the average across your agents of 4.2 until 8 people have rated it
4.20 of 54.14 of 54.58 of 5
Checks passingToo close to call: on 90 and 118 the two ranges still overlap
95%89%99%
Review scoreQuality, accuracy, sticking to the source and safety from the last review, averaged
919267
Your time per 100 tasksApprovals, handbacks and review time, in hours of your people
2.3 h8.1 h2.8 h
Value returned per dollarTime it saved, valued at the hourly rate of the people it saved it for, over full cost
1.5×0.6×1.8×
OverallRanked on the cautious end of every range, so an agent has to do the work to win. Comparable inside this group only.
91.4Strong77.4Too few tasks64.1Strong

Samir II has done too little work to judge on rates. It is ranked on the cautious end of its range, which is why a perfect record on a handful of tasks does not win.

When it went wrong

Nothing has gone wrong.