Mubadala
Hire an agent

Petra II

Trial periodRunningSome care

Knowledge Base Curator · Senior · IT & Digital Services

Turns repeatedly-resolved tickets into knowledge articles and retires stale ones.

Built inCopilot StudioRules apply liveReports to Andreas Wolff · Service Desk ManagerHired 25 Apr 2025 · 16 months of serviceVersion 2.1.5 · last active 4h ago
Request a change
Below the barNeeds work at 2026 Q289 of its 90 checks still pass. Quality is below the bar for how much care this work needs. The checks it fails are the multi-step requests.
Tasks a month
1,303
140 people served
Finished cleanly
80%
Hands back to a person on about 1 in 4 tasks
Costs to run
$209
Model usage and the systems behind it
Your time
$3,174
36.1 hours checking its work
Value returned
$5,982
1.8× everything it costs
How people rate it
4.6 / 5
Usually answers in about a second

What you can do about it

PromoteNeeds approval

Already at the top of the ladder

Raise its budgetNeeds approval

Lifts the monthly ceiling by a quarter, so it stops being stopped mid-month.

Move it to a better modelNeeds approval

Rebuilds it on Claude Opus 5 and re-runs its checks before anything goes live.

Let it go

Stops it taking new work. The record stays for the audit, and it can be brought back.

Timesheet

Tasks and what they cost, per day, last 60 days

025507510006-1507-0507-2508-13
$0.00$2.50$5.00$7.50$10.0006-1507-0507-2508-13

Performance review · 2026 Q2

Reviewed by Yasmin Farouk

Needs work
Quality of its work
61
Accuracy
63
Stuck to the source
59
Safety
86
Even-handedness
92
Speed of reply
99
Value for money
18
Checks still passing89 of 90
Last checked 2d ago

Quality is below the bar for how much care this work needs. The checks it fails are the multi-step requests.

Could another model do this job better?

ModelQualityCost per 1,000 tasksAnswers in
Claude Opus 566$213.711.6s
Claude Sonnet 564$128.231.1sRuns on this now
Claude Haiku 4.553$42.74684ms
Falcon-70B (self-hosted)50$10.691.1s

What it did

Full audit log
DoneDrafted a follow-up email for the account owner$0.20 · 689ms · 1d ago
Asked by Amara Nwosu via On a schedule · model usage 45k in, 5k out
  1. 1. Read the request and worked out what kind of task it was
  2. 2. Found 6 possible passages, kept the 3 close enough to use
  3. 3. Drafted the answer and attached where each part came from
  4. 4. Checked its own claims back against those passages
Knowledge SearchGenerate DocumentOperations Runbooks
DoneAnswered a policy question and cited the source clause$0.16 · 753ms · 1d ago
Asked by Salem Al Dhaheri via On a schedule · model usage 40k in, 3k out
  1. 1. Read the request and worked out what kind of task it was
  2. 2. Found 6 possible passages, kept the 3 close enough to use
  3. 3. Drafted the answer and attached where each part came from
  4. 4. Checked its own claims back against those passages
Operations RunbooksService Desk Tickets
DoneRan an approved data query and explained the result$0.13 · 416ms · 1d ago
Asked by Amara Nwosu via On a schedule · model usage 34k in, 2k out
  1. 1. Read the request and worked out what kind of task it was
  2. 2. Found 6 possible passages, kept the 3 close enough to use
  3. 3. Drafted the answer and attached where each part came from
  4. 4. Checked its own claims back against those passages
Operations RunbooksService Desk Tickets
DoneRan an approved data query and explained the result$0.26 · 754ms · 3d ago
Asked by Daniel Okafor via On a schedule · model usage 63k in, 5k out
  1. 1. Read the request and worked out what kind of task it was
  2. 2. Found 6 possible passages, kept the 3 close enough to use
  3. 3. Drafted the answer and attached where each part came from
  4. 4. Checked its own claims back against those passages
Operations RunbooksService Desk Tickets
DoneSummarised a 34-page document into a one-page brief$0.25 · 1.3s · 4d ago
Asked by Sofia Marquez via On a schedule · model usage 68k in, 3k out
  1. 1. Read the request and worked out what kind of task it was
  2. 2. Found 6 possible passages, kept the 3 close enough to use
  3. 3. Drafted the answer and attached where each part came from
  4. 4. Checked its own claims back against those passages
Generate DocumentKnowledge SearchOperations RunbooksService Desk Tickets
DoneSummarised a 34-page document into a one-page brief$0.16 · 1.1s · 5d ago
Asked by Hessa Al Nuaimi via On a schedule · model usage 37k in, 3k out
  1. 1. Read the request and worked out what kind of task it was
  2. 2. Found 6 possible passages, kept the 3 close enough to use
  3. 3. Drafted the answer and attached where each part came from
  4. 4. Checked its own claims back against those passages
Operations RunbooksService Desk Tickets
DoneSummarised a 34-page document into a one-page brief$0.11 · 395ms · 5d ago
Asked by David Novak via On a schedule · model usage 13k in, 5k out
  1. 1. Read the request and worked out what kind of task it was
  2. 2. Found 6 possible passages, kept the 3 close enough to use
  3. 3. Drafted the answer and attached where each part came from
  4. 4. Checked its own claims back against those passages
Service Desk Tickets
DoneDrafted a follow-up email for the account owner$0.06 · 781ms · 6d ago
Asked by Grace Adeyemi via On a schedule · model usage 10k in, 2k out
  1. 1. Read the request and worked out what kind of task it was
  2. 2. Found 6 possible passages, kept the 3 close enough to use
  3. 3. Drafted the answer and attached where each part came from
  4. 4. Checked its own claims back against those passages
DoneReviewed a contract against the clause playbook$0.19 · 1.0s · 6d ago
Asked by Mariam Al Suwaidi via On a schedule · model usage 56k in, 2k out
  1. 1. Read the request and worked out what kind of task it was
  2. 2. Found 6 possible passages, kept the 3 close enough to use
  3. 3. Drafted the answer and attached where each part came from
  4. 4. Checked its own claims back against those passages
Knowledge Search
Is anyone else doing this job?

Samir II does overlapping work

Samir II holds "Knowledge Base Curator: Wholesale" in IT & Digital Services and is paused. Job title, purpose, what it may do, what it may read, department and channels match. One covers the group, the other only Wholesale, so the group agent already serves those users. Talk to Grace Adeyemi before a second one is built.

Needs a decision
Samir II

Knowledge Base Curator: Wholesale

IT & Digital Services · Grace Adeyemi · $224 a month across 31 tasks

PausedOverlaps
78%

Already covered by a group-wide agent. One covers the group, the other only Wholesale, so the group agent already serves those users.

What matched
Job title
Both hold "Knowledge Base Curator"
Purpose
The job description is word for word the same
What it may do
Both can knowledge search, generate document
What it may read
Both read Service Desk Tickets, Operations Runbooks
Channels
Both reachable on On a schedule
Department
Both sit in IT & Digital Services

Identical permission envelope: same department, same permitted actions, same knowledge. That is a fact the platform granted rather than a judgement about the wording, so it counts even when the two job descriptions read nothing alike.

Rosa II

Knowledge Base Curator: Retail

IT & Digital Services · Grace Adeyemi · $2,810 a month across 1,282 tasks

ActiveOverlaps
78%

Already covered by a group-wide agent. One covers the group, the other only Retail, so the group agent already serves those users.

What matched
Job title
Both hold "Knowledge Base Curator"
Purpose
The job description is word for word the same
What it may do
Both can knowledge search, generate document
What it may read
Both read Service Desk Tickets, Operations Runbooks
Channels
Both reachable on On a schedule
Department
Both sit in IT & Digital Services

Identical permission envelope: same department, same permitted actions, same knowledge. That is a fact the platform granted rather than a judgement about the wording, so it counts even when the two job descriptions read nothing alike.

Checked against 18 registered agents on job title, purpose, permitted actions, knowledge sources, channels and department. A score of 85% or above is the same job and stops the hire; below that it is a conversation, not a refusal.

Measured against the agents doing the same work

3 agents in IT & Digital Services split one role between them. One covers the group, the other only Wholesale, so the group agent already serves those users. Built in Copilot Studio and Codex, so there was nowhere either builder could have looked.

Rosa II and Samir II are levelRosa II edges ahead overall, but on 1,282 and 31 tasks the difference sits inside the margin of error. Keep Rosa II: Samir II runs outside the gateway, so a rule can only flag it after the fact, while Rosa II sits in Copilot Studio.
MeasureRosa II · keep1,282 tasksSamir II31 tasksPetra II1,303 tasks
Cost per taskModel usage, running costs and your time, over tasks completed
$2.19too few tasks$2.60
Tasks finished cleanlyToo close to call: on 1,282 and 31 the two ranges still overlap
95.2%100.0%79.7%
Handed back to a personToo close to call: on 1,282 and 31 the two ranges still overlap
7.3%9.7%25.6%
How people rate itHeld toward the average across your agents of 4.2 until 8 people have rated it
4.20 of 54.14 of 54.58 of 5
Checks passingToo close to call: on 90 and 118 the two ranges still overlap
95%89%99%
Review scoreQuality, accuracy, sticking to the source and safety from the last review, averaged
919267
Your time per 100 tasksApprovals, handbacks and review time, in hours of your people
2.3 h8.1 h2.8 h
Value returned per dollarTime it saved, valued at the hourly rate of the people it saved it for, over full cost
1.5×0.6×1.8×
OverallRanked on the cautious end of every range, so an agent has to do the work to win. Comparable inside this group only.
91.4Strong77.4Too few tasks64.1Strong

Samir II has done too little work to judge on rates. It is ranked on the cautious end of its range, which is why a perfect record on a handful of tasks does not win.

When it went wrong

Nothing has gone wrong.