Caleb II
PausedNot deployedSome careLicence Optimisation Agent: APAC · Junior · IT & Digital Services
Finds unused software licences and proposes reclaim actions to the asset owner.
What you can do about it
Timesheet
Tasks and what they cost, per day, last 60 days
Performance review · 2026 Q2
Reviewed by Yasmin Farouk
Reliable on its main task and says where its answers came from. Answers drift a little on cases it was not built for.
Could another model do this job better?
| Model | Quality | Cost per 1,000 tasks | Answers in | |
|---|---|---|---|---|
| Claude Opus 5 | 99 | $206.67 | 5.3s | |
| Claude Sonnet 5 | 92 | $124.00 | 3.8s | Runs on this now |
| Claude Haiku 4.5 | 89 | $41.33 | 2.3s | |
| Falcon-70B (self-hosted) | 87 | $10.33 | 3.8s |
What it did
DoneGenerated the shift handover pack$0.12 · 1.4s · 4h ago
- 1. Read the request and worked out what kind of task it was
- 2. Found 6 possible passages, kept the 3 close enough to use
- 3. Drafted the answer and attached where each part came from
- 4. Checked its own claims back against those passages
DoneSummarised a 34-page document into a one-page brief$0.32 · 1.7s · 18h ago
- 1. Read the request and worked out what kind of task it was
- 2. Found 6 possible passages, kept the 3 close enough to use
- 3. Drafted the answer and attached where each part came from
- 4. Checked its own claims back against those passages
DoneAnswered a policy question and cited the source clause$0.20 · 2.3s · 19h ago
- 1. Read the request and worked out what kind of task it was
- 2. Found 6 possible passages, kept the 3 close enough to use
- 3. Drafted the answer and attached where each part came from
- 4. Checked its own claims back against those passages
DoneResolved a service desk ticket from the runbook$0.06 · 4.9s · 1d ago
- 1. Read the request and worked out what kind of task it was
- 2. Found 6 possible passages, kept the 3 close enough to use
- 3. Drafted the answer and attached where each part came from
- 4. Checked its own claims back against those passages
Stopped by a ruleTried to open a restricted payroll record$0.26 · 4.7s · 1d ago
- 1. Read the request and worked out what kind of task it was
- 2. Found a document to work from
- 3. A rule stopped it before it read the content
- 4. Stopped the task and recorded it
Waiting for youWaiting on approval to update the CRM record$0.23 · 3.1s · 2d ago
- 1. Read the request and worked out what kind of task it was
- 2. Prepared the change it wanted to make
- 3. A rule caught it: this changes a system that needs approval
- 4. Paused until a named person approves
DoneAnswered a policy question and cited the source clause$0.17 · 2.8s · 3d ago
- 1. Read the request and worked out what kind of task it was
- 2. Found 6 possible passages, kept the 3 close enough to use
- 3. Drafted the answer and attached where each part came from
- 4. Checked its own claims back against those passages
DoneProduced the weekly variance commentary$0.14 · 3.3s · 3d ago
- 1. Read the request and worked out what kind of task it was
- 2. Found 6 possible passages, kept the 3 close enough to use
- 3. Drafted the answer and attached where each part came from
- 4. Checked its own claims back against those passages
DoneReviewed a contract against the clause playbook$0.15 · 2.7s · 4d ago
- 1. Read the request and worked out what kind of task it was
- 2. Found 6 possible passages, kept the 3 close enough to use
- 3. Drafted the answer and attached where each part came from
- 4. Checked its own claims back against those passages
DoneReviewed a contract against the clause playbook$0.08 · 2.3s · 5d ago
- 1. Read the request and worked out what kind of task it was
- 2. Found 6 possible passages, kept the 3 close enough to use
- 3. Drafted the answer and attached where each part came from
- 4. Checked its own claims back against those passages
DoneGenerated the shift handover pack$0.05 · 4.0s · 6d ago
- 1. Read the request and worked out what kind of task it was
- 2. Found 6 possible passages, kept the 3 close enough to use
- 3. Drafted the answer and attached where each part came from
- 4. Checked its own claims back against those passages
FailedThe system it depends on did not answer after three tries$0.05 · 3.9s · 6d ago
- 1. Read the request and worked out what kind of task it was
- 2. Called the system it depends on
- 3. Tried again three times, all three failed
- 4. Reported the failure back to whoever asked
Anton II does overlapping work
Anton II holds "Licence Optimisation Agent" in IT & Digital Services and is doing that work every day. Job title, purpose, what it may do, what it may read, department and channels match. One covers the group, the other only APAC, so the group agent already serves those users. Talk to Yasmin Farouk before a second one is built.
Licence Optimisation Agent
IT & Digital Services · Yasmin Farouk · $602 a month across 830 tasks
Already covered by a group-wide agent. One covers the group, the other only APAC, so the group agent already serves those users.
What matched
- Job title
- Both hold "Licence Optimisation Agent"
- Purpose
- The job description is word for word the same
- What it may do
- Both can warehouse query, post to teams
- What it may read
- Both read Enterprise Data Warehouse, Vendor & Service Provider Data
- Channels
- Both reachable on On a schedule
- Department
- Both sit in IT & Digital Services
Identical permission envelope: same department, same permitted actions, same knowledge. That is a fact the platform granted rather than a judgement about the wording, so it counts even when the two job descriptions read nothing alike.
Licence Optimisation Agent: Field
IT & Digital Services · Yasmin Farouk · $537 a month across 864 tasks
Same role, split by scope. Scoped to APAC and Field respectively.
What matched
- Job title
- Both hold "Licence Optimisation Agent"
- Purpose
- The job description is word for word the same
- What it may do
- Both can warehouse query, post to teams
- What it may read
- Both read Enterprise Data Warehouse, Vendor & Service Provider Data
- Channels
- Both reachable on On a schedule
- Department
- Both sit in IT & Digital Services
Identical permission envelope: same department, same permitted actions, same knowledge. That is a fact the platform granted rather than a judgement about the wording, so it counts even when the two job descriptions read nothing alike.
Measured against the agents doing the same work
3 agents in IT & Digital Services split one role between them. One covers the group, the other only Field, so the group agent already serves those users. Built in Azure AI Foundry and Microsoft 365 Copilot, so there was nowhere either builder could have looked.
| Measure | Zara II · keep864 tasks | Anton II830 tasks | Caleb II90 tasks |
|---|---|---|---|
Cost per taskModel usage, running costs and your time, over tasks completed | $0.62● | $0.72 | $5.87 |
Tasks finished cleanlyToo close to call: on 90 and 830 the two ranges still overlap | 94.0% | 94.6% | 100.0% |
Handed back to a personToo close to call: on 830 and 864 the two ranges still overlap | 6.5% | 6.4% | 11.1% |
How people rate itToo close to call: on 48 and 4 the two ranges still overlap | 3.75 of 5 | 4.11 of 5 | 4.04 of 5 |
Checks passingToo close to call: on 188 and 153 the two ranges still overlap | 87% | 93% | 89% |
Review scoreQuality, accuracy, sticking to the source and safety from the last review, averaged | 97● | 86 | 95 |
Your time per 100 tasksApprovals, handbacks and review time, in hours of your people | 0.7 h● | 0.7 h | 6.6 h |
Value returned per dollarTime it saved, valued at the hourly rate of the people it saved it for, over full cost | 7.6×● | 1.7× | 0.8× |
OverallRanked on the cautious end of every range, so an agent has to do the work to win. Comparable inside this group only. | 89.7Moderate | 89.6Moderate | 71.4Thin |
Caleb II has done too little work to judge on rates. It is ranked on the cautious end of its range, which is why a perfect record on a handful of tasks does not win.