# Why Asama Is Different: Infrastructure State and the Detect, Investigate, Remediate Loop

> What it takes to run the detect, investigate, remediate loop on a large fleet, why stale and stitched data breaks it, and why Asama is built on infrastructure state.

- URL: https://blog.asama.ai/blog/why-asama-is-different/
- Published: 2026-09-21
- Tags: Articles

## Executive summary

Running the detect, investigate, remediate loop well on a fleet of a few thousand servers takes far more than alerts and dashboards. Every step consumes specific inputs: logs and counters, but also each node's firmware versions, BIOS attributes, kernel settings and tuning, the topology that connects the nodes, the intent they are measured against, the change history that moved them and the procedures that fix them. Each input goes stale on its own clock, and a stale one doesn't stay in its stage. Its error multiplies through the rest of the loop and shortens the useful life of every diagnosis and runbook built on it.

Most teams keep those inputs in five or more tools that can't agree on what a server is called, and a senior engineer joins them by hand.

The missing layer is infrastructure state: a live, normalised, vendor-neutral record of what every component is, beyond what it is doing. Monitoring collects behaviour, configuration management enforces the settings you chose to manage, and the CMDB stores what you meant to deploy. None of them holds the state that explains behaviour, which is why bolting a language model onto a monitoring stack improves its correlation and barely moves its coverage.

Asama is built from that layer up: in-band and out-of-band collection, normalisation into one object per physical part, a live fleet model, deterministic correlation before any agentic reasoning, and remediation planned against the exact state of each box. On five issues run side by side at a design partner, Asama reached root cause in 3 minutes 20 seconds on average. By hand: 65 minutes.

## What the loop needs from every node

The loop is really two loops. Detection runs from symptom through correlation and root cause to verification; remediation runs from plan through action to verification; both should learn. Correlation is the pivot of the first, because that is where the hours go. Once the right signals are tied together, one or two candidate causes remain and naming the right one is quick. Planning is the pivot of the second. Running a command is easy. Knowing which command is right for this box, on today's firmware, inside this week's change policy, is the hard part.

Take a representative incident. A chassis temperature alert fires on a compute node during a heavy load run, and the dashboard agrees the node is hot. Stand in the cold aisle. Nothing feels different. If you could pick one node's fans out of the roar of the hall, they would sound a little quieter than their neighbours' under the same load, and that is the only physical tell.

To get from that alert to a verified fix, you need to establish that inlet temperature and power draw are normal, so it isn't the room or the supply; that CPU package temperature is climbing and clocks are throttling; and that fan RPM is low against identical peers. You need the BMC audit log entry showing the fan profile moving from performance to optimised six days ago, right after a BMC firmware update, and the list of every other node on that build. Before touching anything, you also need the cohort's approved fan profile, the next change window and how many nodes can be out at once.

That is nine facts from at least six sources, for one alert.

Across every fault class and every layer, the inventory looks like this.

| Input | What it answers | Examples on a real node | Stages that need it |
|---|---|---|---|
| Events and logs | What happened? | BMC SEL and audit log; kernel ring buffer (MCE, EDAC, PCIe AER, NVMe resets); journald; GPU XID; hypervisor and kubelet events | Symptom, correlation |
| Metrics and counters | How is it behaving? | CPU steal and softirq; NIC drop and softnet counters; NVMe latency; SMART and NVMe health log; inlet and exhaust temperature; fan RPM; PSU input; PCIe correctable-error and replay counters; DCGM fields | Symptom, correlation, verification |
| Inventory and identity | What parts are in the box, and which record is which part? | SKU and serials; DIMM, drive, NIC, HBA and GPU per slot; PCIe addresses; the mapping between a block device, a controller bay, a Redfish resource and a caddy serial | Correlation, blast radius, action |
| Versions | What does it run? | BIOS/UEFI, BMC, CPLD, NIC, HBA, drive and GPU firmware; CPU microcode; kernel; drivers; package set | State detection, root cause, plan |
| Configuration | How is it set? | BIOS attributes (power profile, C-states, SMT, ASPM, SR-IOV); BMC settings (fan profile, power cap); kernel command line; sysctl; loaded modules; enabled services | Drift detection, root cause, plan |
| Tuning | How was it hand-tuned? | IRQ affinity; CPU governor; NUMA placement and vCPU pinning; hugepages; NIC ring sizes and RSS queues; MTU | Root cause, plan |
| Topology | What is it connected to, and what state is each link in? | VM to host; pod to node; host to leaf port; NUMA and PCIe tree; HBA to array; PSU to PDU feed; rack; cohort | Correlation, blast radius, wave planning |
| Intent | What should it be? | Golden config per cohort; vendor compatibility matrix; CVE and errata feeds; learned baselines per workload regime | State detection, verification |
| Change and policy | What moved, and what may move? | Change requests; firmware jobs; package updates; maintenance windows; freezes; PodDisruptionBudgets; max-unavailable | Root cause, plan, action |
| History | How did this box and its cohort get here? | Past incidents and their verified causes; RMAs; prior fixes and whether they held | Root cause, plan, learning |
| Remediation knowledge | How do you fix it safely on this exact box? | Procedures with preconditions; pre-computed rollback; verification checks; drain and cordon steps; canary sizing | Plan, action, verification |

Monitoring stacks are built to collect the first two rows. Root causes and safe fixes live in the other nine.

## Every input has a TTL

Every row in that table decays on its own clock. A fan RPM reading is good for seconds. A BIOS attribute holds until the next firmware job, the next RMA or the next time someone clears CMOS. An IRQ affinity mask lasts until the next reboot, or until a package update quietly enables irqbalance (the whole point of a meta-package is that nobody reads what it pulls in). A runbook lasts until a new SKU, firmware build or kernel lands on the nodes it targets.

Here is where each input usually lives in a stitched-together stack, and what breaks when the copy you're reasoning over is older than the box.

| Input | Changes when | Usually lives in | What breaks when it's stale |
|---|---|---|---|
| Sensor and counter readings | Every second | TSDB, BMC | Poll every minute and a thermal spike that clears in 30 seconds never happened. |
| Firmware and BIOS state | Every BMC job, BIOS update, reflash or RMA | Vendor console, spreadsheet, CMDB | Root cause misses the change. The fix targets the wrong build. |
| Kernel, drivers, packages, sysctl | Every patch run or image rebuild | Config-management facts from the last run | Drift reads as a mystery. Playbook preconditions check a world that no longer exists. |
| Tuning | Reboots, daemons, package defaults | Usually nowhere | Performance anomalies with no recorded cause |
| Inventory and identity | Every part swap, RMA or re-rack | CMDB, asset database | The wrong drive gets pulled from a degraded array. The wrong node gets drained. |
| Topology | Live migrations in minutes, recabling in months | Hypervisor manager, cabling sheet, DCIM | Blast radius is wrong in both directions. |
| Baselines | Workload regime shifts, firmware changes, new SKUs | Thresholds somebody set years ago | False positives your team learns to ignore, false negatives nobody sees |
| Intent and advisories | Every vendor advisory and policy change | Wiki, vendor portals | Drift and exposure go undetected, because nothing current exists to compare against. |
| Change and policy | Every CR, freeze and window | ITSM | A fix lands in a freeze, or a wave exceeds max-unavailable. |
| History | Every incident | Tickets, postmortems, people | The same fault is investigated from zero. Lemon nodes cycle back into the pool. |
| Runbooks and playbooks | Whenever someone rewrites them | Wiki, Git | The wrong procedure runs against the wrong hardware state. |

If I could fix one row first, it would be inventory and identity, because every other row joins on it.

The last row is the one you can measure. Asama's fleet observations put about 18 percent of manual fixes at the wrong procedure for the node's actual hardware state. Usually the procedure was right when someone wrote it. The box changed under it.

### Errors multiply down the loop

Stages don't average their errors. They multiply them. If symptom detection, correlation, root cause, plan and verification are each right 90 percent of the time, the loop produces a correct, verified fix 59 percent of the time; let one stale input drag a single stage to 70 percent and the loop falls to 46 percent. A downstream stage can't correct an upstream error it has no independent data to check against, so it treats whatever it receives as fact.

A stale value is worse than a missing one. A missing BIOS attribute makes a system say it doesn't know. A stale one makes it wrong with confidence, and it hands that confidence to the next stage.

### The half-life of a diagnosis

Everything the loop produces, whether a finding, a root cause, a baseline or a runbook, is computed from state and stays useful until that state changes. So a diagnosis lives only as long as the stalest input it touched, whatever the average age of the rest. A root cause computed from a firmware inventory last reconciled a month ago can't be trusted on any node that has taken a firmware job since. A runbook written for last quarter's SKU mix starts decaying the day the next procurement batch is racked.

Stale inputs shorten the life of every artefact built on them. Your team compensates the only way it can: by rebuilding context by hand, every incident.

### Precise enough to act on

Accuracy asks whether an answer is right. Precision asks whether it is specific enough to act on. "CPU pressure on the host" can be accurate and still useless to whoever holds the pager. "BMC firmware build N reset the fan profile to optimised on the fourteen nodes that took it six days ago, and those nodes throttle under sustained load" names the attribute, the build and the node set, which is exactly what a change plan needs.

No system is more precise than its inputs. If the attribute, the build and the node set aren't held somewhere as current, comparable state, the most precise honest answer on offer is the directional one.

## Five tool types, four handoffs, and you as the join

Here is the fan-profile incident in a typical stack.

1. Zabbix or Prometheus fires on a static chassis temperature threshold.
2. You open Grafana. The node is hot. Inlet temperature looks normal, assuming someone enabled IPMI polling, which Zabbix ships switched off.
3. You open iDRAC or iLO for fan RPM, then a second node's console to find out what normal RPM looks like.
4. You SSH in and run turbostat. Package temperature is up and clocks are down.
5. Back in the BMC audit log, you find the fan profile change.
6. In ServiceNow or Jira you find a BMC firmware job from six days ago.
7. OpenManage, OneView or a spreadsheet tells you which other nodes took the same build, if it's current.
8. The wiki has a runbook for the previous hardware generation, so you adapt a Redfish script.
9. You raise a change request, wait for the change board, and roll the fix in waves inside the next window, checking fans node by node.
10. The finding ends up in a ticket comment, where the next incident won't look.

Each of those tools is internally consistent. Together they are not. Monitoring knows the node by hostname, the vendor console by service tag, the BMC by IP address, the CMDB by a CI name that was right at install time and Ansible by an inventory alias; each refreshes on its own schedule and keeps its own clock, so no two of them agree on which drive sits in bay 3.

You are the join.

Asama estimates that tools do about 30 percent of today's detection work, 15 percent of investigation and 25 percent of remediation, and people do the rest. On the fleets Asama has measured, firefighting of this kind takes 40 to 60 percent of senior engineering time. None of that cost lives inside one tool. It sits at the intersections, which is why a better version of any single tool barely moves MTTR. Capacity pays too: 7 to 15 percent of a typical fleet sits in the maintenance queue at any moment, depreciating while it waits on this manual loop.

## Infrastructure state, defined

Infrastructure state is the complete, current, normalised record of what every component in your estate is: its identity, its parts, its firmware and software versions, its configuration and tuning, its place in the topology, what it is supposed to be, and how it came to be that way.

It sits next to three other things that are easy to blur together.

| | Question it answers | In the fan-profile incident | Who usually holds it |
|---|---|---|---|
| Behaviour | What is the box doing? | The CPU is hot and throttling | Monitoring |
| State | What is the box? | Fan profile set to optimised; BMC on build N | Nobody, consistently |
| Intent | What should the box be? | This cohort runs the performance profile | A wiki page, a golden image, someone's memory |
| Change and history | How did it get here, and what happened last time? | The BMC job six days ago; whether an earlier firmware cycle did the same | Tickets, logs, people |

Behaviour is a function of state, workload and environment. Monitoring watches the output of that function and tries to infer the inputs, which is an underdetermined problem, because many different states produce the same hot, throttling CPU. When a model built on behaviour alone leaves a large residual with structure in it, that residual is usually a state variable nobody collected.

### Half the fault tree is state

Split infrastructure faults cleanly and you get two kinds. Runtime anomalies are problems of behaviour, covering availability, correctness, performance, saturation and activity. State deviations are problems of what the box is: configuration drift from golden config, version and lifecycle skew against the fleet and the vendor matrix, a component set that doesn't match approved inventory, and known defects from CVE and errata feeds.

A tool that doesn't hold state cannot detect the second kind at all.

A BIOS setting reset to a vendor default emits no metric. Nothing fires. Days or weeks later it surfaces as a runtime anomaly, and by then the tool sees the anomaly and not the cause, which is exactly how the fan-profile incident played out six days after the firmware job.

### What holding state takes

Five things, and each has a failure mode your senior engineer currently covers by hand.

| Component | What it means | What a person does today when it's missing |
|---|---|---|
| Breadth | Every layer is collected: components and interconnect, firmware and BMC, BIOS, drivers, kernel and OS, up to the hypervisor and Kubernetes, across vendors | SSHes into the box, logs into the BMC |
| Depth | Past events and telemetry to configuration and tuning state, judged against a reference that evolves with the fleet | Remembers what normal looks like |
| Normalisation | Every record means the same thing across vendors and layers, and one physical part is one object | Knows both vendors' dialects, matches serials by eye |
| Topology | A live graph of what connects to what, with state on the edges | Holds the estate's wiring in their head |
| History | Prior faults, RMAs and changes for the node and its cohort, surfaced into the diagnosis unasked | Remembers what happened in March |

Asama's [published evaluation](https://blog.asama.ai/blog/01-asama-vs-zabbixclaude) calls these five together Infrastructure Coverage, and scores every platform on them alongside detection and remediation. Coverage caps everything above it. A reasoning layer can't correlate a layer nobody collected, and it can't plan a change against firmware it has never seen.

## Why your tools don't hold it

Nobody decided infrastructure state wasn't worth collecting. It fell between the categories the market was built around.

**Monitoring is built on the time series.** A TSDB stores a float against a timestamp. A BIOS attribute registry with a few hundred vendor-specific enumerations, an IRQ affinity mask, a PCIe tree and a firmware tuple are documents and graphs. Pressed into labels or text items, they can be displayed but not compared.

**Every vendor speaks a dialect.** Redfish standardised the transport and the shape of the schema, and every OEM still ships its own attribute registry and its own Oem extensions. OEM SEL records need vendor-specific decoding. Firmware version strings don't compare across vendors, and sometimes not across generations from the same vendor. Keeping all of that current is an engineering bill that comes due again with every hardware generation, per vendor. It never ends.

**Vendor consoles see only their own boxes.** OpenManage knows Dell, OneView knows HPE and XClarity knows Lenovo, each in depth, and each knows next to nothing about the machine from another vendor in the next rack. Customers rarely let one OEM's tooling manage another OEM's servers, so the party that best understands a box's state can't compare it with its peers.

**Cloud-native tooling assumed hardware was somebody else's job.** The cloud-era observability platforms grew up on hyperscaler infrastructure, where the provider hides the BIOS, the BMC and the firmware from you. Teams running their own metal inherited tools designed for a world where that layer wasn't theirs.

**The CMDB records intent, and config management enforces a subset.** Ansible, Puppet and Salt keep the parameters you chose to manage consistent, and they're very good at it. Firmware defaults, BIOS profiles, driver behaviour and any tuning you never wrote into a role sit outside their view, and their facts are a snapshot from the last run. The CMDB describes what you meant to build.

**Nobody owns the firmware layer.** Ask who owns the application, the network or the OS and you get a team name. Ask who owns BIOS settings, BMC firmware, microcode and driver baselines and the room goes quiet.

**The LLM wave built reasoning, not collection.** Wire a frontier model to Zabbix through its API and you get conversational correlation over everything Zabbix holds, in an afternoon. In [Asama's published scoring](https://blog.asama.ai/blog/01-asama-vs-zabbixclaude), that pairing lifts Zabbix's detection score from 0.8 to 2.3 out of 4 and its Infrastructure Coverage from 1.4 to 1.8, all of the coverage gain coming from searchable history. A model reading a store adds no layer, no vendor and no data class to it. Across the 23 platforms and combinations in that scoring, L3 is common on correlation and root cause, and rare or absent on breadth, depth, normalisation and topology. The framework is Asama's own. Every anchor is published so you can re-score any cell, and it has already been corrected once, after an early version counted the same scoped logs as both correlation and verification and inflated five well-known platforms by a full level.

The operators who do hold infrastructure state built it for themselves. [Meta's research clusters](https://arxiv.org/abs/2410.21680) run systematic lemon-node detection, and taking those nodes out of scheduling cut the failure rate of 512-plus-GPU jobs from 14 percent to 4 percent. [Crusoe](https://www.crusoe.ai/resources/blog/autoclusters-minimizing-hardware-failures-in-large-gpu-clusters), [Nebius](https://nebius.com/blog/posts/how-we-build-reliable-clusters) and Together each engineered multi-source hardware health checks in-house. That took platform teams and years, and none of it is for sale.

## What Asama built, from the BMC up

Asama is an AI infrastructure engineer for teams that run their own servers. It isn't a monitoring tool, a dashboard or a chat window over your alerts: it holds the state those tools don't, then does the investigation and planning your engineers do now. It was built in the opposite order from most of the market, with collection and the data model first and reasoning last.

### Collection: in-band and out-of-band

A host agent reads the kernel, eBPF probes, /proc and /sys, and the system logs. The BMC is read out-of-band over Redfish and IPMI for the SEL, sensors, firmware versions, BIOS attributes and audit log, so a box can still be interrogated when its OS can't answer. IPMI, which most BMCs still answer on UDP port 623, dates back to 1998. GPUs report through DCGM and XID events, and drives through SMART and the NVMe health log. Kubernetes, OpenStack, Proxmox and Slurm supply the platform layer, while CMDB entries, change requests and golden configs come in as intent and change. Critical signals are sampled every second.

### Normalisation: one part, one object

This is the expensive layer. Vendor knowledge decodes every raw SEL, XID, AER and SMART code into a part, a fault and a severity. A Dell power reading and an HPE power reading become the same quantity. The block device the kernel enumerates, the bay the controller reports, the Redfish drive resource and the serial on the caddy label resolve to one object with one identity, so a finding can name the physical part a technician has to pull.

In one recorded session, the engineer asked Asama about packet drops and named the wrong host. Asama worked out which machine was actually dropping before it started the investigation.

### The fleet model

On top of normalised records, Asama keeps one live model of the fleet (topology, baselines, configuration and history) and updates it as the fleet changes. Every baseline is versioned. Each version is tied to the event that triggered it and to the peer group in force at the time, so any past detection can be audited later. Peer groups are explicit and built on SKU, firmware, workload regime and age together. Two nodes of the same model on different firmware are not peers. Firmware rollouts, part swaps and other known fleet events are registered, and a step change that lines up with one is read as a legitimate break rather than an anomaly. The model also tracks its own sensor coverage and says so when a failure class is only partly visible on a node. [Normal Is a Moving Target](https://blog.asama.ai/blog/normal-is-a-moving-target) sets out the statistics behind this.

### Detection: four comparisons before anything fires

Every node is compared four ways: against intent (golden config), against advisories (CVE and errata), against its peers and against its own learned baseline. Before a finding fires, it is checked for fleet context (one box or many?), workload regime (normal for this load?) and sensor coverage (can this reading be trusted?). What reaches you is a finding with a cause attached. Peer comparison is also what catches the drive that has run three times slower than its neighbours since the day it was racked; Asama's fleet observations put about 65 servers in every 1,000 in that silently degraded state, where no threshold sees them. Across 50-plus induced failures on Dell, HPE and Lenovo hardware, detection took 8 to 50 seconds.

### Investigation: deterministic first, agentic last

Correlation runs on four dimensions: time (aligned sequences across sources), topology (what shares this host's path), cohort (do identical peers show the same symptom?) and state and change (what moved, and when). About 70 percent of Asama's root-cause work is deterministic correlation and statistical inference. The agentic 30 percent takes those results plus the raw evidence and writes the causal chain, so the model never guesses at a fact the fleet model already holds. When it can't read something, it says so.

In a recorded CPU-steal investigation on a lab hypervisor, the guest showed steal while its host was nowhere near busy. Asama found that the guest's vCPUs and the host NIC's interrupts were confined to the same four host cores, so receive processing competed with the guest for CPU time. It ruled out host saturation, cited the IRQ affinity it read from /proc/irq, and reported that it could not read irqbalance's status rather than guessing at it. Time to root cause: 2 minutes 50 seconds.

Every RCA arrives with the chain, the evidence, a confidence score with its reasons, and the blast radius: the other nodes carrying the same at-risk state right now.

### Remediation: planned against the box you actually have

A plan is composed at incident time for the node's exact hardware, firmware, OS and configuration, with preconditions and a rollback worked out before anything runs. It is then fitted to your policy (maintenance windows, freezes, PodDisruptionBudgets, max-unavailable), and the change request is raised for you. Execution goes dry run, canary, then waves. Verification compares each node with its own pre-incident baseline and with its peers, so a fix counts only when the cause is gone. The outcome updates the procedure library, and the next plan starts from what actually worked on your fleet.

Autonomy is earned in stages. On day one your engineers approve every plan. With trust established, they pre-approve mitigations such as cordon, drain and failover, and once a fix class has a track record, they pre-approve it in policy: a firmware rollback, say, or a config revert.

### Execution: SaaS plans, your premises execute

The reasoning engine runs in Asama's SOC 2 SaaS and never touches your infrastructure. It emits a signed, tamper-evident manifest bound to intent and scope. On your premises the executor verifies the signature with keys you hold, rejects anything unsigned or out of scope, and runs deterministic steps with least privilege. BMC credentials stay in your own Vault. The collector needs 2 vCPU, 8 GB of RAM and 50 GB of disk, and connects outbound on 443 only. Air-gapped and fully on-prem control planes are available for regulated environments. Every action is logged and replayable.

Nothing gets ripped out: Prometheus, Grafana, Datadog and Elastic feed Asama as signal sources, ServiceNow and Jira receive incidents with the RCA attached, and PagerDuty, Opsgenie, Slack and Teams carry the pages and the approvals.

## Why this team could build it

You can't get here by [extending an L3 stack](https://blog.asama.ai/blog/why_l3_stalls). Better reasoning over the same inputs pays off fast and then flattens, and the work that actually moves the constraint, such as collecting BIOS attributes across five OEMs or resolving four records of one drive into one object, doesn't improve the product a vendor is already shipping. It makes a different product possible, and very few vendors can justify the detour. Ask any vendor, this one included, which half they built first. In Asama's [four-level maturity model](https://blog.asama.ai/blog/four-levels-of-detection-maturity), that detour is the gap between L3, where the platform assists and your engineer still drives the diagnosis, and L4, where the platform holds a live model of the stack and your engineer approves.

Asama started there.

Its founders ran infrastructure at Flipkart: more than 35,000 servers across three data centres of 4 to 6 MW each, India's largest GPU fleet, 3.6 million telemetry events a second and 20 PB of data processed a day, on over $500M of infrastructure spend across ten years. Sumeet, the CEO, was Head of Infrastructure. JJ, the CTO, was Lead Architect for cloud orchestration and platform services. At that scale they were the human join, and they know which vendor quirks matter because each one cost them incident hours. The ten-person engineering team spans systems engineering, AI and ML research.

Asama is also vendor-neutral by construction. One model covers Dell, HPE, Lenovo, Supermicro and NVIDIA hardware, which no OEM console can offer.

## The fan-profile incident, rerun on Asama

The same alert, on Asama:

- **Finding.** Fourteen nodes in one cohort throttling under load, raised as one finding.
- **Chain.** BMC firmware build N, applied six days ago, reset the fan profile from performance to optimised. Under load the fans run slower than on cohort peers, CPU package temperature climbs, clocks throttle, and the chassis alert fires.
- **Evidence.** The BMC audit log entry, fan RPM against peers under matching load and the firmware job in the change record, with inlet temperature and power ruled out. Confidence scored, with reasons.
- **Blast radius.** Every node that took build N, including three that haven't run a heavy job since and so haven't paged anyone.
- **Plan.** Restore the fan profile over Redfish on every affected node, with a canary on two and then waves in the next window. Pin the attribute in the cohort's golden config, so any node that takes build N later is flagged as drift on the day it happens.
- **Verification.** Fan RPM back inside the cohort band and clocks holding under sustained load, measured against each node's own pre-incident baseline and against its peers, before the incident is closed and the fix is written back into the procedure library.

Your engineer reads the chain and approves the plan.

## What it did on a design partner's fleet

In a pilot with a mid-market design partner that runs its own data centres, Asama produced 26 root-cause analyses. Twenty were accepted on first read. Three needed one follow-up question before they were right. The other three were incomplete, because part of the cause sat outside what the platform covers, and the team finished the correlation by hand. On five issues worked side by side, the team's own engineers averaged 65 minutes to root cause per issue, while Asama, working the same issues from the same alerts, averaged 3 minutes 20 seconds.

Those three incomplete RCAs mark where the product is today. Visibility is generally available, detection is validated, and RCA and remediation are expanding. Network reasoning is the active build.

One question is still open, for Asama as much as anyone: whether partial state coverage buys you anything before all five components are in place. Asama's view is that it buys close to nothing, because a peer comparison over unnormalised data compares a Dell reading with an HPE one and means nothing. Nobody has proven it.

## Run the test on your own fleet

Ask your current stack five questions.

1. Which nodes are running a BIOS attribute that differs from the rest of their cohort, right now?
2. What caused your last incident, and which other nodes carry that cause today?
3. Which physical drive is throwing errors, by bay and serial, from one query?
4. Did your last automated fix check the node's firmware build before it ran?
5. What did your platform learn from your last fix?

I'd start with the first one. Most stacks can't answer it at all.

If a person answers instead of the platform, the gap is infrastructure state, and a better model on top won't close it.

Asama fits teams that run their own servers on-prem, in colo or in a private cloud, from a few hundred nodes up, on mixed hardware, with a small infra team carrying the pager. It deploys in under two hours. The pilot runs 60 days, with visibility on day one, RCA on your own incidents by day 15, customisation to your stack through day 45, and a joint MTTD and MTTR review against your baseline at the end. Run it read-only beside your current tools at first, and count what it finds that they didn't.

Write to contactus@asama.ai, or book a live comparison against your own stack at [asama.ai](https://www.asama.ai).
