Mubadala
Hire an agent

Tariq II

Being builtRunningHigh care

Device Health Monitor · Senior · IT & Digital Services

Spots devices about to fail from the status they report and raises replacement tickets early.

Built inCopilot StudioRules apply liveReports to Yasmin Farouk · Head of ITHired 3 Mar 2025 · 18 months of serviceVersion 3.6.2 · last active 15h ago
Request a change
Acting alone on careful workActs on its own on high care workIt changes things in other systems with nobody checking first, and asked for approval only 13 times in 30 days. Confirm that is deliberate and on the record.
Below the barOn an improvement plan at 2026 Q2128 of its 163 checks still pass. Failing its checks and handing back work a person then has to redo. Retire it, or rebuild it from the current template.
Two agents, one jobSame job as Elena IITariq II and Elena II both hold "Device Health Monitor" in IT & Digital Services. Merging them leaves one set of instructions to keep up to date and one record to answer for. Compare with Elena II
Costs more than it returnsReturns 0.5× what it costsModel usage and running costs are only 2% of the bill: the rest is 16 hours a month of your people approving, unpicking and checking its work. Below 1.5× it is not paying for the time it takes up.
Tasks a month
371
16 people served
Finished cleanly
72%
Hands back to a person on about 1 in 4 tasks
Costs to run
$29
Model usage and the systems behind it
Your time
$1,404
16 hours checking its work
Value returned
$716
0.5× everything it costs
How people rate it
4.1 / 5
Usually answers in about 3 seconds

What you can do about it

PromoteNeeds approval

Already at the top of the ladder

Raise its budgetNeeds approval

Lifts the monthly ceiling by a quarter, so it stops being stopped mid-month.

Move it to a better modelNeeds approval

Rebuilds it on Claude Opus 5 and re-runs its checks before anything goes live.

Let it go

Stops it taking new work. The record stays for the audit, and it can be brought back.

Timesheet

Tasks and what they cost, per day, last 60 days

01325385006-1507-0507-2508-13
$0.00$0.50$1.00$1.50$2.0006-1507-0507-2508-13

Performance review · 2026 Q2

Reviewed by Yasmin Farouk

On an improvement plan
Quality of its work
55
Accuracy
54
Stuck to the source
51
Safety
84
Even-handedness
85
Speed of reply
89
Value for money
18
Checks still passing128 of 163
Last checked 13d ago

Failing its checks and handing back work a person then has to redo. Retire it, or rebuild it from the current template.

Could another model do this job better?

ModelQualityCost per 1,000 tasksAnswers in
Claude Opus 562$114.913.2s
Claude Sonnet 556$68.952.3sRuns on this now
Claude Haiku 4.552$22.981.4s
Falcon-70B (self-hosted)49$5.752.3s

What it did

Full audit log
DoneDrafted a follow-up email for the account owner$0.14 · 2.5s · 1d ago
Asked by Claire Dubois via On a schedule · model usage 45k in, 538 out
  1. 1. Read the request and worked out what kind of task it was
  2. 2. Found 6 possible passages, kept the 3 close enough to use
  3. 3. Drafted the answer and attached where each part came from
  4. 4. Checked its own claims back against those passages
Warehouse QueryCreate Ticket
DoneRan an approved data query and explained the result$0.17 · 4.1s · 1d ago
Asked by Layla Al Mazrouei via On a schedule · model usage 47k in, 2k out
  1. 1. Read the request and worked out what kind of task it was
  2. 2. Found 6 possible passages, kept the 3 close enough to use
  3. 3. Drafted the answer and attached where each part came from
  4. 4. Checked its own claims back against those passages
Create TicketWarehouse Query
DoneProduced the weekly variance commentary$0.22 · 3.1s · 2d ago
Asked by Mariam Al Suwaidi via On a schedule · model usage 52k in, 4k out
  1. 1. Read the request and worked out what kind of task it was
  2. 2. Found 6 possible passages, kept the 3 close enough to use
  3. 3. Drafted the answer and attached where each part came from
  4. 4. Checked its own claims back against those passages
Warehouse Query
DoneGenerated the shift handover pack$0.32 · 2.6s · 4d ago
Asked by Daniel Okafor via On a schedule · model usage 86k in, 4k out
  1. 1. Read the request and worked out what kind of task it was
  2. 2. Found 6 possible passages, kept the 3 close enough to use
  3. 3. Drafted the answer and attached where each part came from
  4. 4. Checked its own claims back against those passages
Create TicketService Desk Tickets
DoneAnswered a policy question and cited the source clause$0.11 · 601ms · 4d ago
Asked by Claire Dubois via On a schedule · model usage 30k in, 2k out
  1. 1. Read the request and worked out what kind of task it was
  2. 2. Found 6 possible passages, kept the 3 close enough to use
  3. 3. Drafted the answer and attached where each part came from
  4. 4. Checked its own claims back against those passages
Create TicketWarehouse Query
DoneSummarised a 34-page document into a one-page brief$0.17 · 2.4s · 5d ago
Asked by Layla Al Mazrouei via On a schedule · model usage 46k in, 2k out
  1. 1. Read the request and worked out what kind of task it was
  2. 2. Found 6 possible passages, kept the 3 close enough to use
  3. 3. Drafted the answer and attached where each part came from
  4. 4. Checked its own claims back against those passages
DoneRan an approved data query and explained the result$0.18 · 2.1s · 5d ago
Asked by Layla Al Mazrouei via On a schedule · model usage 54k in, 1k out
  1. 1. Read the request and worked out what kind of task it was
  2. 2. Found 6 possible passages, kept the 3 close enough to use
  3. 3. Drafted the answer and attached where each part came from
  4. 4. Checked its own claims back against those passages
Create TicketWarehouse QueryService Desk Tickets
Stopped by a ruleThe model it runs on is not on the approved list$0.30 · 2.5s · 5d ago
Asked by Claire Dubois via On a schedule · model usage 82k in, 3k out
  1. 1. Read the request and worked out what kind of task it was
  2. 2. Found a document to work from
  3. 3. A rule stopped it before it read the content
  4. 4. Stopped the task and recorded it
Create TicketService Desk Tickets
Stopped by a rule: Personal data never leaves
Is anyone else doing this job?

Elena II already does this job

Elena II holds "Device Health Monitor" in IT & Digital Services and is doing that work every day. Job title, purpose, what it may do, what it may read, department and channels match. Extend that agent, or change this description until the two jobs are genuinely different.

Blocked
Elena II

Device Health Monitor

IT & Digital Services · Grace Adeyemi · $779 a month across 205 tasks

ActiveSame job
100%

Same job: job title, purpose, what it may do all match. Neither title narrows the job to a region or a business line, so both serve the same people.

What matched
Job title
Both hold "Device Health Monitor"
Purpose
The job description is word for word the same
What it may do
Both can warehouse query, create ticket
What it may read
Both read Enterprise Data Warehouse, Service Desk Tickets
Channels
Both reachable on On a schedule
Department
Both sit in IT & Digital Services

Identical permission envelope: same department, same permitted actions, same knowledge. That is a fact the platform granted rather than a judgement about the wording, so it counts even when the two job descriptions read nothing alike.

Checked against 18 registered agents on job title, purpose, permitted actions, knowledge sources, channels and department. A score of 85% or above is the same job and stops the hire; below that it is a conversation, not a refusal.

Measured against the agents doing the same work

2 agents in IT & Digital Services hold the same job. The pair is an outright copy, not a split by region or business line. Built in Azure AI Foundry and Copilot Studio, so there was nowhere either builder could have looked.

Elena II is the one to keepElena II scores 34.9 points ahead of Tariq II, completing 100.0% of 205 tasks against 72.0% of 371 at $3.80 against $3.86 a task.
MeasureElena II · keep205 tasksTariq II371 tasks
Cost per taskModel usage, running costs and your time, over tasks completed
$3.80$3.86
Tasks finished cleanlyRange allows for how few tasks some agents have run
100.0%72.0%
Handed back to a personEvery handback is a task the agent did not finish
11.2%26.4%
How people rate itToo close to call: on 37 and 16 the two ranges still overlap
4.11 of 54.12 of 5
Checks passingShare of its checks that passed at the last run
90%79%
Review scoreQuality, accuracy, sticking to the source and safety from the last review, averaged
9361
Your time per 100 tasksApprovals, handbacks and review time, in hours of your people
4.3 h4.3 h
Value returned per dollarTime it saved, valued at the hourly rate of the people it saved it for, over full cost
0.5×0.5×
OverallRanked on the cautious end of every range, so an agent has to do the work to win. Comparable inside this group only.
90.8Moderate55.9Moderate

When it went wrong

MediumDepartment went over budgetOpen1mo ago

The department passed its monthly budget. Because work stops at the limit, further tasks were refused until the budget was raised.

HighDepartment went over budgetClosed1mo ago

The department passed its monthly budget. Because work stops at the limit, further tasks were refused until the budget was raised.