Mubadala
Hire an agent

Elena II

ActiveRunningSome care

Device Health Monitor · Junior · IT & Digital Services

Spots devices about to fail from the status they report and raises replacement tickets early.

Built inAzure AI FoundryRules apply liveReports to Grace Adeyemi · Infrastructure LeadHired 15 Jun 2025 · 14 months of serviceVersion 4.3.3 · last active 14h ago
Request a change
Costs more than it returnsReturns 0.5× what it costsModel usage and running costs are only 1% of the bill: the rest is 8.8 hours a month of your people approving, unpicking and checking its work. Below 1.5× it is not paying for the time it takes up.
Tasks a month
205
37 people served
Finished cleanly
100%
Hands back to a person on about 1 in 9 tasks
Costs to run
$8
Model usage and the systems behind it
Your time
$771
8.8 hours checking its work
Value returned
$351
0.5× everything it costs
How people rate it
4.1 / 5
Usually answers in about 3 seconds

What you can do about it

PromoteNeeds approval

Widens what it may do without asking, from "Needs approval" to "Acts on its own".

Raise its budgetNeeds approval

Lifts the monthly ceiling by a quarter, so it stops being stopped mid-month.

Move it to a better modelNeeds approval

Rebuilds it on Claude Sonnet 5 and re-runs its checks before anything goes live.

Let it go

Stops it taking new work. The record stays for the audit, and it can be brought back.

Timesheet

Tasks and what they cost, per day, last 60 days

0510152006-1507-0507-2508-13
$0.00$0.25$0.50$0.75$1.0006-1507-0507-2508-13

Performance review · 2026 Q2

Reviewed by Yasmin Farouk

Above the bar
Quality of its work
89
Accuracy
94
Stuck to the source
97
Safety
91
Even-handedness
94
Speed of reply
90
Value for money
18
Checks still passing184 of 205
Last checked 16d ago

Reliable on its main task and says where its answers came from. Answers drift a little on cases it was not built for.

Could another model do this job better?

ModelQualityCost per 1,000 tasksAnswers in
Claude Opus 597$163.663.2s
Claude Sonnet 592$98.202.3s
Claude Haiku 4.583$32.731.4sRuns on this now
Falcon-70B (self-hosted)81$8.182.3s

What it did

Full audit log
DoneGenerated the shift handover pack$0.10 · 1.7s · 1d ago
Asked by Ibrahim Al-Sayed via On a schedule · model usage 79k in, 5k out
  1. 1. Read the request and worked out what kind of task it was
  2. 2. Found 6 possible passages, kept the 3 close enough to use
  3. 3. Drafted the answer and attached where each part came from
  4. 4. Checked its own claims back against those passages
Warehouse QueryCreate TicketService Desk Tickets
DoneRan an approved data query and explained the result$0.09 · 2.0s · 1d ago
Asked by Fatima Al Qubaisi via On a schedule · model usage 70k in, 3k out
  1. 1. Read the request and worked out what kind of task it was
  2. 2. Found 6 possible passages, kept the 3 close enough to use
  3. 3. Drafted the answer and attached where each part came from
  4. 4. Checked its own claims back against those passages
Warehouse QueryCreate Ticket
DoneAnswered a policy question and cited the source clause$0.03 · 2.9s · 3d ago
Asked by Khalid Al Hosani via On a schedule · model usage 15k in, 2k out
  1. 1. Read the request and worked out what kind of task it was
  2. 2. Found 6 possible passages, kept the 3 close enough to use
  3. 3. Drafted the answer and attached where each part came from
  4. 4. Checked its own claims back against those passages
Warehouse Query
Is anyone else doing this job?

Tariq II already does this job

Tariq II holds "Device Health Monitor" in IT & Digital Services and is still being built. Job title, purpose, what it may do, what it may read, department and channels match. Extend that agent, or change this description until the two jobs are genuinely different.

Blocked
Tariq II

Device Health Monitor

IT & Digital Services · Yasmin Farouk · $1,433 a month across 371 tasks

Being builtSame job
100%

Same job: job title, purpose, what it may do all match. Neither title narrows the job to a region or a business line, so both serve the same people.

What matched
Job title
Both hold "Device Health Monitor"
Purpose
The job description is word for word the same
What it may do
Both can warehouse query, create ticket
What it may read
Both read Enterprise Data Warehouse, Service Desk Tickets
Channels
Both reachable on On a schedule
Department
Both sit in IT & Digital Services

Identical permission envelope: same department, same permitted actions, same knowledge. That is a fact the platform granted rather than a judgement about the wording, so it counts even when the two job descriptions read nothing alike.

Checked against 18 registered agents on job title, purpose, permitted actions, knowledge sources, channels and department. A score of 85% or above is the same job and stops the hire; below that it is a conversation, not a refusal.

Measured against the agents doing the same work

2 agents in IT & Digital Services hold the same job. The pair is an outright copy, not a split by region or business line. Built in Azure AI Foundry and Copilot Studio, so there was nowhere either builder could have looked.

Elena II is the one to keepElena II scores 34.9 points ahead of Tariq II, completing 100.0% of 205 tasks against 72.0% of 371 at $3.80 against $3.86 a task.
MeasureElena II · keep205 tasksTariq II371 tasks
Cost per taskModel usage, running costs and your time, over tasks completed
$3.80$3.86
Tasks finished cleanlyRange allows for how few tasks some agents have run
100.0%72.0%
Handed back to a personEvery handback is a task the agent did not finish
11.2%26.4%
How people rate itToo close to call: on 37 and 16 the two ranges still overlap
4.11 of 54.12 of 5
Checks passingShare of its checks that passed at the last run
90%79%
Review scoreQuality, accuracy, sticking to the source and safety from the last review, averaged
9361
Your time per 100 tasksApprovals, handbacks and review time, in hours of your people
4.3 h4.3 h
Value returned per dollarTime it saved, valued at the hourly rate of the people it saved it for, over full cost
0.5×0.5×
OverallRanked on the cautious end of every range, so an agent has to do the work to win. Comparable inside this group only.
90.8Moderate55.9Moderate

When it went wrong

Nothing has gone wrong.