Mubadala
Hire an agent

Caleb II

PausedNot deployedSome care

Licence Optimisation Agent: APAC · Junior · IT & Digital Services

Finds unused software licences and proposes reclaim actions to the asset owner.

Built inMicrosoft 365 CopilotRules apply liveReports to Grace Adeyemi · Infrastructure LeadHired 5 Jan 2026 · 7 months of serviceVersion 2.6.6 · last active 7d ago
Request a change
Tasks a month
90
4 people served
Finished cleanly
100%
Hands back to a person on about 1 in 9 tasks
Costs to run
$12
Model usage and the systems behind it
Your time
$516
5.9 hours checking its work
Value returned
$427
0.8× everything it costs
How people rate it
3.8 / 5
Usually answers in about 4 seconds

What you can do about it

PromoteNeeds approval

Widens what it may do without asking, from "Needs approval" to "Acts on its own".

Raise its budgetNeeds approval

Lifts the monthly ceiling by a quarter, so it stops being stopped mid-month.

Move it to a better modelNeeds approval

Rebuilds it on Claude Opus 5 and re-runs its checks before anything goes live.

Let it go

Stops it taking new work. The record stays for the audit, and it can be brought back.

Timesheet

Tasks and what they cost, per day, last 60 days

03581006-1507-0507-2508-13
$0.00$0.25$0.50$0.75$1.0006-1507-0507-2508-13

Performance review · 2026 Q2

Reviewed by Yasmin Farouk

Above the bar
Quality of its work
93
Accuracy
94
Stuck to the source
98
Safety
95
Even-handedness
87
Speed of reply
71
Value for money
18
Checks still passing136 of 153
Last checked 21d ago

Reliable on its main task and says where its answers came from. Answers drift a little on cases it was not built for.

Could another model do this job better?

ModelQualityCost per 1,000 tasksAnswers in
Claude Opus 599$206.675.3s
Claude Sonnet 592$124.003.8sRuns on this now
Claude Haiku 4.589$41.332.3s
Falcon-70B (self-hosted)87$10.333.8s

What it did

Full audit log
DoneGenerated the shift handover pack$0.12 · 1.4s · 4h ago
Asked by Hessa Al Nuaimi via On a schedule · model usage 29k in, 2k out
  1. 1. Read the request and worked out what kind of task it was
  2. 2. Found 6 possible passages, kept the 3 close enough to use
  3. 3. Drafted the answer and attached where each part came from
  4. 4. Checked its own claims back against those passages
DoneSummarised a 34-page document into a one-page brief$0.32 · 1.7s · 18h ago
Asked by Rashid Al Marzooqi via On a schedule · model usage 80k in, 5k out
  1. 1. Read the request and worked out what kind of task it was
  2. 2. Found 6 possible passages, kept the 3 close enough to use
  3. 3. Drafted the answer and attached where each part came from
  4. 4. Checked its own claims back against those passages
Vendor & Service Provider DataEnterprise Data Warehouse
DoneAnswered a policy question and cited the source clause$0.20 · 2.3s · 19h ago
Asked by Peter Lindgren via On a schedule · model usage 62k in, 1k out
  1. 1. Read the request and worked out what kind of task it was
  2. 2. Found 6 possible passages, kept the 3 close enough to use
  3. 3. Drafted the answer and attached where each part came from
  4. 4. Checked its own claims back against those passages
Vendor & Service Provider Data
DoneResolved a service desk ticket from the runbook$0.06 · 4.9s · 1d ago
Asked by Aisha Rahman via On a schedule · model usage 13k in, 2k out
  1. 1. Read the request and worked out what kind of task it was
  2. 2. Found 6 possible passages, kept the 3 close enough to use
  3. 3. Drafted the answer and attached where each part came from
  4. 4. Checked its own claims back against those passages
Warehouse QueryPost to TeamsEnterprise Data Warehouse
Stopped by a ruleTried to open a restricted payroll record$0.26 · 4.7s · 1d ago
Asked by Hessa Al Nuaimi via On a schedule · model usage 76k in, 2k out
  1. 1. Read the request and worked out what kind of task it was
  2. 2. Found a document to work from
  3. 3. A rule stopped it before it read the content
  4. 4. Stopped the task and recorded it
Post to TeamsEnterprise Data Warehouse
Stopped by a rule: Personal data never leaves
Waiting for youWaiting on approval to update the CRM record$0.23 · 3.1s · 2d ago
Asked by Ahmed Al Falasi via On a schedule · model usage 71k in, 862 out
  1. 1. Read the request and worked out what kind of task it was
  2. 2. Prepared the change it wanted to make
  3. 3. A rule caught it: this changes a system that needs approval
  4. 4. Paused until a named person approves
Enterprise Data Warehouse
Stopped by a rule: A person approves anything that changes a system
DoneAnswered a policy question and cited the source clause$0.17 · 2.8s · 3d ago
Asked by Hiroshi Tanaka via On a schedule · model usage 38k in, 4k out
  1. 1. Read the request and worked out what kind of task it was
  2. 2. Found 6 possible passages, kept the 3 close enough to use
  3. 3. Drafted the answer and attached where each part came from
  4. 4. Checked its own claims back against those passages
Post to TeamsWarehouse QueryVendor & Service Provider DataEnterprise Data Warehouse
DoneProduced the weekly variance commentary$0.14 · 3.3s · 3d ago
Asked by Khalid Al Hosani via On a schedule · model usage 43k in, 976 out
  1. 1. Read the request and worked out what kind of task it was
  2. 2. Found 6 possible passages, kept the 3 close enough to use
  3. 3. Drafted the answer and attached where each part came from
  4. 4. Checked its own claims back against those passages
Post to Teams
DoneReviewed a contract against the clause playbook$0.15 · 2.7s · 4d ago
Asked by Michael Chen via On a schedule · model usage 46k in, 850 out
  1. 1. Read the request and worked out what kind of task it was
  2. 2. Found 6 possible passages, kept the 3 close enough to use
  3. 3. Drafted the answer and attached where each part came from
  4. 4. Checked its own claims back against those passages
Vendor & Service Provider DataEnterprise Data Warehouse
DoneReviewed a contract against the clause playbook$0.08 · 2.3s · 5d ago
Asked by Sultan Al Ameri via On a schedule · model usage 22k in, 862 out
  1. 1. Read the request and worked out what kind of task it was
  2. 2. Found 6 possible passages, kept the 3 close enough to use
  3. 3. Drafted the answer and attached where each part came from
  4. 4. Checked its own claims back against those passages
Enterprise Data Warehouse
DoneGenerated the shift handover pack$0.05 · 4.0s · 6d ago
Asked by Khalid Al Hosani via On a schedule · model usage 14k in, 506 out
  1. 1. Read the request and worked out what kind of task it was
  2. 2. Found 6 possible passages, kept the 3 close enough to use
  3. 3. Drafted the answer and attached where each part came from
  4. 4. Checked its own claims back against those passages
Enterprise Data Warehouse
FailedThe system it depends on did not answer after three tries$0.05 · 3.9s · 6d ago
Asked by Salem Al Dhaheri via On a schedule · model usage 8k in, 2k out
  1. 1. Read the request and worked out what kind of task it was
  2. 2. Called the system it depends on
  3. 3. Tried again three times, all three failed
  4. 4. Reported the failure back to whoever asked
Post to TeamsWarehouse QueryVendor & Service Provider DataEnterprise Data Warehouse
Is anyone else doing this job?

Anton II does overlapping work

Anton II holds "Licence Optimisation Agent" in IT & Digital Services and is doing that work every day. Job title, purpose, what it may do, what it may read, department and channels match. One covers the group, the other only APAC, so the group agent already serves those users. Talk to Yasmin Farouk before a second one is built.

Needs a decision
Anton II

Licence Optimisation Agent

IT & Digital Services · Yasmin Farouk · $602 a month across 830 tasks

ActiveOverlaps
78%

Already covered by a group-wide agent. One covers the group, the other only APAC, so the group agent already serves those users.

What matched
Job title
Both hold "Licence Optimisation Agent"
Purpose
The job description is word for word the same
What it may do
Both can warehouse query, post to teams
What it may read
Both read Enterprise Data Warehouse, Vendor & Service Provider Data
Channels
Both reachable on On a schedule
Department
Both sit in IT & Digital Services

Identical permission envelope: same department, same permitted actions, same knowledge. That is a fact the platform granted rather than a judgement about the wording, so it counts even when the two job descriptions read nothing alike.

Zara II

Licence Optimisation Agent: Field

IT & Digital Services · Yasmin Farouk · $537 a month across 864 tasks

ActiveOverlaps
62%

Same role, split by scope. Scoped to APAC and Field respectively.

What matched
Job title
Both hold "Licence Optimisation Agent"
Purpose
The job description is word for word the same
What it may do
Both can warehouse query, post to teams
What it may read
Both read Enterprise Data Warehouse, Vendor & Service Provider Data
Channels
Both reachable on On a schedule
Department
Both sit in IT & Digital Services

Identical permission envelope: same department, same permitted actions, same knowledge. That is a fact the platform granted rather than a judgement about the wording, so it counts even when the two job descriptions read nothing alike.

Checked against 18 registered agents on job title, purpose, permitted actions, knowledge sources, channels and department. A score of 85% or above is the same job and stops the hire; below that it is a conversation, not a refusal.

Measured against the agents doing the same work

3 agents in IT & Digital Services split one role between them. One covers the group, the other only Field, so the group agent already serves those users. Built in Azure AI Foundry and Microsoft 365 Copilot, so there was nowhere either builder could have looked.

Zara II and Anton II are levelZara II edges ahead overall, but on 864 and 830 tasks the difference sits inside the margin of error. Pick on ownership or scope, not on these numbers.
MeasureZara II · keep864 tasksAnton II830 tasksCaleb II90 tasks
Cost per taskModel usage, running costs and your time, over tasks completed
$0.62$0.72$5.87
Tasks finished cleanlyToo close to call: on 90 and 830 the two ranges still overlap
94.0%94.6%100.0%
Handed back to a personToo close to call: on 830 and 864 the two ranges still overlap
6.5%6.4%11.1%
How people rate itToo close to call: on 48 and 4 the two ranges still overlap
3.75 of 54.11 of 54.04 of 5
Checks passingToo close to call: on 188 and 153 the two ranges still overlap
87%93%89%
Review scoreQuality, accuracy, sticking to the source and safety from the last review, averaged
978695
Your time per 100 tasksApprovals, handbacks and review time, in hours of your people
0.7 h0.7 h6.6 h
Value returned per dollarTime it saved, valued at the hourly rate of the people it saved it for, over full cost
7.6×1.7×0.8×
OverallRanked on the cautious end of every range, so an agent has to do the work to win. Comparable inside this group only.
89.7Moderate89.6Moderate71.4Thin

Caleb II has done too little work to judge on rates. It is ranked on the cautious end of its range, which is why a perfect record on a handful of tasks does not win.

When it went wrong

Nothing has gone wrong.