Frequently asked questions

Questions we get
Measurement, reports and assurance

What triage actually is, what a cluster assessment produces, and how the findings report, the delta report and the residual-risk report differ. If your question is not here, just ask it.

What we do

Triage, assessment and assurance

What does triage mean in this context?

Triage comes from emergency care: sort by severity first, treat afterwards. Around a cluster incident it means the signals are ordered before anyone starts changing things. The free triage script on ClusterDown reads the event channels of all nodes across the window you specify, and returns the five heaviest signals, sorted by severity, each with the measured value, the node and the timestamp.

Triage sorts, it does not explain. Establishing the cause is the step after that.

↑ Back to the questions

What is a cluster assessment?

A full measurement of cluster, hosts and the dependencies around them, followed by assessment by a specialist. A collection script runs from your own management server over WinRM to the nodes. It only reads and changes nothing; the base measurement takes about fifteen minutes.

What comes out is not a dashboard score but a report in which every finding carries the measured value, the norm, the deviation, the risk, the repair and the check afterwards. How that path runs is set out on the method page.

↑ Back to the questions

What is Assurance?

Assurance is the cluster assessment as an annual cycle instead of a single measurement: one baseline (C1) and three follow-up measurements (V1, V2 and V3), each using the same measurement logic and each followed by a report and a one-hour review.

That makes visible whether remediation actually worked and what configuration drift appeared between rounds. A clean report in January says nothing about October. The full arrangement is on the Assurance page.

↑ Back to the questions

How does ClusterDown differ from ClusterTriage Assurance?

ClusterDown is the free entry point after an incident: one read-only script, the five heaviest signals, no report and no subscription. Assurance is the annual cycle for an environment you want to keep demonstrably under control: four full measurements, four reports and four reviews.

The same specialists and the same read-only method, a different purpose.

↑ Back to the questions

The reports

What you get back, and how the reports differ

What is the difference between the Findings, Delta and Residual-risk report?

The three reports come from the same measurement and answer three different questions.

  • Findings report. The complete, prioritised technical picture of this round. Every finding carries the measured value, the norm with its source, the deviation, the risk, the repair and the check with which you confirm the result.
  • Delta report. The comparison with the previous round: what was resolved, what is still open and what is new. This is the verification. A closed ticket does not prove the environment changed, a follow-up measurement does.
  • Residual-risk report. What remains open after the repair work, with its consequence. Not the original problem list, but the current risk position.

Alongside those, the residual-risk trend across four rounds shows whether risk is falling, holding steady or climbing again.

↑ Back to the questions

Which report do I get when?

The baseline produces the findings report. From the first follow-up measurement onwards, each round adds the delta report and the current residual risk. After four rounds you hold four comparable measurement points and a trend line across the year.

↑ Back to the questions

What exactly is in a single finding?

Six fixed parts: what was measured, with values, node names and date; what the vendor or Microsoft prescribes, with the source; where the two diverge and what that means in your situation; the repair work in order, with the commands; the check with which you confirm it worked, the same command we found it with; and how you keep the result in place.

What was found to be in order is in the report too. That is your evidence of what was examined.

↑ Back to the questions

Which languages is the report available in?

Dutch, English and German. All languages are built from the same measurement and are not translated afterwards.

↑ Back to the questions

The measurement in practice

What happens on your systems, and what does not

Does the script change anything in our environment?

No. It reads, and nothing else. No agent, no service and no scheduled task; nothing is left behind afterwards. It changes no settings and restarts nothing, and it works with the privileges of the administrator who starts it.

Before the file is written, the script checks its own output for values that look like credentials.

↑ Back to the questions

How long does it take, and what do you need from us?

The base measurement takes about fifteen minutes. On top of that a round costs you an hour for the review. What is needed is a management server that can reach the nodes, and a one-off intake: name, nodes, subnet and domain, plus a conversation about what we may see.

From that follows which modules make sense and which access they require; we request that access up front. No support contract, Azure subscription or workspace is required.

↑ Back to the questions

What happens to the measurement data?

The measurement produces one file on your own server. You place it on the encrypted share we set up for you, and we use it solely to produce your report. In the free triage on ClusterDown, memory dumps stay out of scope.

↑ Back to the questions

And if you cannot measure something?

Then it is named in the report, with the reason: no access to the array, management controllers unreachable, or a node that did not respond. So you know what the report covers and what it does not. What we could not see does not quietly disappear from the conclusion.

↑ Back to the questions

Working together

What we do for you, and what we do not

Do you carry out the repair work as well?

The repair is yours to do. The report is written so your own team or your supplier can act on it directly: the steps are in order, with the commands and with the check afterwards. On a time-and-materials basis we can also do it for you.

↑ Back to the questions

Is this monitoring or operations?

Neither. We monitor nothing, install nothing and do not take over operations. This is a measurement at an agreed moment, not alerting. Your existing monitoring stays exactly where it is.

The assessment lays side by side the layers that monitoring usually sees separately: host, cluster, storage, network, array, management controllers, Azure and backup. A path that is right on the host and wrong on the switch only stands out when someone compares both.

↑ Back to the questions

How do we start?

With a conversation about your environment and about what we may see. Then you run the measurement, on your own or with us alongside, place the result on the encrypted share, and we deliver the report together with a review session.

If you first want to know what a recent incident has to tell you, start with the free triage on ClusterDown. For an engagement, arrange a conversation.

↑ Back to the questions

ClusterTriage

Is your question not here?

Ask it in a short conversation. We would rather answer concretely than generally.