Tariq II
Being builtRunningHigh careDevice Health Monitor · Senior · IT & Digital Services
Spots devices about to fail from the status they report and raises replacement tickets early.
What you can do about it
Timesheet
Tasks and what they cost, per day, last 60 days
Performance review · 2026 Q2
Reviewed by Yasmin Farouk
Failing its checks and handing back work a person then has to redo. Retire it, or rebuild it from the current template.
Could another model do this job better?
| Model | Quality | Cost per 1,000 tasks | Answers in | |
|---|---|---|---|---|
| Claude Opus 5 | 62 | $114.91 | 3.2s | |
| Claude Sonnet 5 | 56 | $68.95 | 2.3s | Runs on this now |
| Claude Haiku 4.5 | 52 | $22.98 | 1.4s | |
| Falcon-70B (self-hosted) | 49 | $5.75 | 2.3s |
What it did
DoneDrafted a follow-up email for the account owner$0.14 · 2.5s · 1d ago
- 1. Read the request and worked out what kind of task it was
- 2. Found 6 possible passages, kept the 3 close enough to use
- 3. Drafted the answer and attached where each part came from
- 4. Checked its own claims back against those passages
DoneRan an approved data query and explained the result$0.17 · 4.1s · 1d ago
- 1. Read the request and worked out what kind of task it was
- 2. Found 6 possible passages, kept the 3 close enough to use
- 3. Drafted the answer and attached where each part came from
- 4. Checked its own claims back against those passages
DoneProduced the weekly variance commentary$0.22 · 3.1s · 2d ago
- 1. Read the request and worked out what kind of task it was
- 2. Found 6 possible passages, kept the 3 close enough to use
- 3. Drafted the answer and attached where each part came from
- 4. Checked its own claims back against those passages
DoneGenerated the shift handover pack$0.32 · 2.6s · 4d ago
- 1. Read the request and worked out what kind of task it was
- 2. Found 6 possible passages, kept the 3 close enough to use
- 3. Drafted the answer and attached where each part came from
- 4. Checked its own claims back against those passages
DoneAnswered a policy question and cited the source clause$0.11 · 601ms · 4d ago
- 1. Read the request and worked out what kind of task it was
- 2. Found 6 possible passages, kept the 3 close enough to use
- 3. Drafted the answer and attached where each part came from
- 4. Checked its own claims back against those passages
DoneSummarised a 34-page document into a one-page brief$0.17 · 2.4s · 5d ago
- 1. Read the request and worked out what kind of task it was
- 2. Found 6 possible passages, kept the 3 close enough to use
- 3. Drafted the answer and attached where each part came from
- 4. Checked its own claims back against those passages
DoneRan an approved data query and explained the result$0.18 · 2.1s · 5d ago
- 1. Read the request and worked out what kind of task it was
- 2. Found 6 possible passages, kept the 3 close enough to use
- 3. Drafted the answer and attached where each part came from
- 4. Checked its own claims back against those passages
Stopped by a ruleThe model it runs on is not on the approved list$0.30 · 2.5s · 5d ago
- 1. Read the request and worked out what kind of task it was
- 2. Found a document to work from
- 3. A rule stopped it before it read the content
- 4. Stopped the task and recorded it
Elena II already does this job
Elena II holds "Device Health Monitor" in IT & Digital Services and is doing that work every day. Job title, purpose, what it may do, what it may read, department and channels match. Extend that agent, or change this description until the two jobs are genuinely different.
Same job: job title, purpose, what it may do all match. Neither title narrows the job to a region or a business line, so both serve the same people.
What matched
- Job title
- Both hold "Device Health Monitor"
- Purpose
- The job description is word for word the same
- What it may do
- Both can warehouse query, create ticket
- What it may read
- Both read Enterprise Data Warehouse, Service Desk Tickets
- Channels
- Both reachable on On a schedule
- Department
- Both sit in IT & Digital Services
Identical permission envelope: same department, same permitted actions, same knowledge. That is a fact the platform granted rather than a judgement about the wording, so it counts even when the two job descriptions read nothing alike.
Measured against the agents doing the same work
2 agents in IT & Digital Services hold the same job. The pair is an outright copy, not a split by region or business line. Built in Azure AI Foundry and Copilot Studio, so there was nowhere either builder could have looked.
| Measure | Elena II · keep205 tasks | Tariq II371 tasks |
|---|---|---|
Cost per taskModel usage, running costs and your time, over tasks completed | $3.80● | $3.86 |
Tasks finished cleanlyRange allows for how few tasks some agents have run | 100.0%● | 72.0% |
Handed back to a personEvery handback is a task the agent did not finish | 11.2%● | 26.4% |
How people rate itToo close to call: on 37 and 16 the two ranges still overlap | 4.11 of 5 | 4.12 of 5 |
Checks passingShare of its checks that passed at the last run | 90%● | 79% |
Review scoreQuality, accuracy, sticking to the source and safety from the last review, averaged | 93● | 61 |
Your time per 100 tasksApprovals, handbacks and review time, in hours of your people | 4.3 h● | 4.3 h |
Value returned per dollarTime it saved, valued at the hourly rate of the people it saved it for, over full cost | 0.5× | 0.5×● |
OverallRanked on the cautious end of every range, so an agent has to do the work to win. Comparable inside this group only. | 90.8Moderate | 55.9Moderate |
When it went wrong
The department passed its monthly budget. Because work stops at the limit, further tasks were refused until the budget was raised.
The department passed its monthly budget. Because work stops at the limit, further tasks were refused until the budget was raised.