Mubadala
Hire an agent

Rosa II

ActiveRunningRoutine

Knowledge Base Curator: Retail · Junior · IT & Digital Services

Turns repeatedly-resolved tickets into knowledge articles and retires stale ones.

Built inCopilot StudioRules apply liveReports to Grace Adeyemi · Infrastructure LeadHired 28 Dec 2025 · 8 months of serviceVersion 2.8.11 · last active 10h ago
Request a change
Tasks a month
1,282
69 people served
Finished cleanly
95%
Hands back to a person on about 1 in 14 tasks
Costs to run
$167
Model usage and the systems behind it
Your time
$2,643
30 hours checking its work
Value returned
$4,224
1.5× everything it costs
How people rate it
4.2 / 5
Usually answers in about 5 seconds

What you can do about it

PromoteNeeds approval

Widens what it may do without asking, from "Needs approval" to "Acts on its own".

Raise its budgetNeeds approval

Lifts the monthly ceiling by a quarter, so it stops being stopped mid-month.

Move it to a better modelNeeds approval

Rebuilds it on Claude Opus 5 and re-runs its checks before anything goes live.

Let it go

Stops it taking new work. The record stays for the audit, and it can be brought back.

Timesheet

Tasks and what they cost, per day, last 60 days

025507510006-1507-0507-2508-13
$0.00$2.50$5.00$7.50$10.0006-1507-0507-2508-13

Performance review · 2026 Q2

Reviewed by Yasmin Farouk

Meets the bar
Quality of its work
92
Accuracy
91
Stuck to the source
84
Safety
95
Even-handedness
85
Speed of reply
63
Value for money
18
Checks still passing112 of 118
Last checked 11d ago

Does the job it was given. Sticks to the source less well on long documents, worth narrowing what it reads.

Could another model do this job better?

ModelQualityCost per 1,000 tasksAnswers in
Claude Opus 597$189.306.1s
Claude Sonnet 596$113.584.4sRuns on this now
Claude Haiku 4.584$37.862.7s
Falcon-70B (self-hosted)83$9.474.4s

What it did

Full audit log
DoneSummarised a 34-page document into a one-page brief$0.26 · 1.9s · 1d ago
Asked by Ibrahim Al-Sayed via On a schedule · model usage 80k in, 1k out
  1. 1. Read the request and worked out what kind of task it was
  2. 2. Found 6 possible passages, kept the 3 close enough to use
  3. 3. Drafted the answer and attached where each part came from
  4. 4. Checked its own claims back against those passages
Knowledge SearchService Desk TicketsOperations Runbooks
FailedFound no documents it is allowed to read$0.14 · 4.6s · 2d ago
Asked by Nadia Haddad via On a schedule · model usage 27k in, 4k out
  1. 1. Read the request and worked out what kind of task it was
  2. 2. Called the system it depends on
  3. 3. Tried again three times, all three failed
  4. 4. Reported the failure back to whoever asked
DoneSummarised a 34-page document into a one-page brief$0.20 · 2.1s · 3d ago
Asked by Khalid Al Hosani via On a schedule · model usage 43k in, 4k out
  1. 1. Read the request and worked out what kind of task it was
  2. 2. Found 6 possible passages, kept the 3 close enough to use
  3. 3. Drafted the answer and attached where each part came from
  4. 4. Checked its own claims back against those passages
DoneRan an approved data query and explained the result$0.11 · 738ms · 4d ago
Asked by Noura Al Kaabi via On a schedule · model usage 13k in, 5k out
  1. 1. Read the request and worked out what kind of task it was
  2. 2. Found 6 possible passages, kept the 3 close enough to use
  3. 3. Drafted the answer and attached where each part came from
  4. 4. Checked its own claims back against those passages
Generate DocumentKnowledge SearchService Desk TicketsOperations Runbooks
DoneGenerated the shift handover pack$0.09 · 4.4s · 6d ago
Asked by Omar Al Shamsi via On a schedule · model usage 12k in, 3k out
  1. 1. Read the request and worked out what kind of task it was
  2. 2. Found 6 possible passages, kept the 3 close enough to use
  3. 3. Drafted the answer and attached where each part came from
  4. 4. Checked its own claims back against those passages
Generate DocumentKnowledge Search
Is anyone else doing this job?

Petra II does overlapping work

Petra II holds "Knowledge Base Curator" in IT & Digital Services and is doing the work on a trial. Job title, purpose, what it may do, what it may read, department and channels match. One covers the group, the other only Retail, so the group agent already serves those users. Talk to Andreas Wolff before a second one is built.

Needs a decision
Petra II

Knowledge Base Curator

IT & Digital Services · Andreas Wolff · $3,383 a month across 1,303 tasks

Trial periodOverlaps
78%

Already covered by a group-wide agent. One covers the group, the other only Retail, so the group agent already serves those users.

What matched
Job title
Both hold "Knowledge Base Curator"
Purpose
The job description is word for word the same
What it may do
Both can knowledge search, generate document
What it may read
Both read Service Desk Tickets, Operations Runbooks
Channels
Both reachable on On a schedule
Department
Both sit in IT & Digital Services

Identical permission envelope: same department, same permitted actions, same knowledge. That is a fact the platform granted rather than a judgement about the wording, so it counts even when the two job descriptions read nothing alike.

Samir II

Knowledge Base Curator: Wholesale

IT & Digital Services · Grace Adeyemi · $224 a month across 31 tasks

PausedOverlaps
62%

Same role, split by scope. Scoped to Retail and Wholesale respectively.

What matched
Job title
Both hold "Knowledge Base Curator"
Purpose
The job description is word for word the same
What it may do
Both can knowledge search, generate document
What it may read
Both read Service Desk Tickets, Operations Runbooks
Channels
Both reachable on On a schedule
Department
Both sit in IT & Digital Services

Identical permission envelope: same department, same permitted actions, same knowledge. That is a fact the platform granted rather than a judgement about the wording, so it counts even when the two job descriptions read nothing alike.

Checked against 18 registered agents on job title, purpose, permitted actions, knowledge sources, channels and department. A score of 85% or above is the same job and stops the hire; below that it is a conversation, not a refusal.

Measured against the agents doing the same work

3 agents in IT & Digital Services split one role between them. One covers the group, the other only Wholesale, so the group agent already serves those users. Built in Copilot Studio and Codex, so there was nowhere either builder could have looked.

Rosa II and Samir II are levelRosa II edges ahead overall, but on 1,282 and 31 tasks the difference sits inside the margin of error. Keep Rosa II: Samir II runs outside the gateway, so a rule can only flag it after the fact, while Rosa II sits in Copilot Studio.
MeasureRosa II · keep1,282 tasksSamir II31 tasksPetra II1,303 tasks
Cost per taskModel usage, running costs and your time, over tasks completed
$2.19too few tasks$2.60
Tasks finished cleanlyToo close to call: on 1,282 and 31 the two ranges still overlap
95.2%100.0%79.7%
Handed back to a personToo close to call: on 1,282 and 31 the two ranges still overlap
7.3%9.7%25.6%
How people rate itHeld toward the average across your agents of 4.2 until 8 people have rated it
4.20 of 54.14 of 54.58 of 5
Checks passingToo close to call: on 90 and 118 the two ranges still overlap
95%89%99%
Review scoreQuality, accuracy, sticking to the source and safety from the last review, averaged
919267
Your time per 100 tasksApprovals, handbacks and review time, in hours of your people
2.3 h8.1 h2.8 h
Value returned per dollarTime it saved, valued at the hourly rate of the people it saved it for, over full cost
1.5×0.6×1.8×
OverallRanked on the cautious end of every range, so an agent has to do the work to win. Comparable inside this group only.
91.4Strong77.4Too few tasks64.1Strong

Samir II has done too little work to judge on rates. It is ranked on the cautious end of its range, which is why a perfect record on a handful of tasks does not win.

When it went wrong

Nothing has gone wrong.