novasean
← All articles

Operations and recovery

When a client website fails: build the responsibility map before the incident

Use a symptom-to-owner map to coordinate DNS, hosting, runtime, application and supplier investigations without leaving an agency client between teams.

Editorial cover with five blank rounded rectangles linked by lines.
On this article

A website can look like one service to its client while depending on a registrar, DNS provider, content-delivery layer, hosting platform, operating system, runtime, database, application and several external integrations.

When the site fails, those layers do not investigate themselves. Without a responsibility map, each supplier can report that its own component looks healthy while the agency still has a broken client journey.

The map does not need to predict every failure. It needs to answer four operational questions:

  1. Who coordinates the incident?
  2. Who can investigate and change each layer?
  3. What evidence moves the issue from one owner to another?
  4. Who tells the client what is known, unknown and happening next?

Start with the symptom, not the assumed cause

“The server is down” is often an interpretation. Record the observable symptom first:

  • the affected hostname and a sanitised URL path, without query strings, tokens, credentials or unnecessary personal data in the shared incident record;
  • the time and timezone;
  • who reported it and from where;
  • the error message or response code;
  • whether the problem affects every user or a subset;
  • which business function is affected;
  • the last known successful check;
  • recent changes, if known.

The NCSC's small-business response guidance recommends gathering what was reported, which services are affected, when the problem appeared, whether customers have noticed and what the potential business impact is. This initial record gives the next team something testable instead of an unsupported diagnosis.

If the full request is genuinely needed for investigation or preservation, keep the unaltered original separately in an authorised, access-controlled evidence store; do not copy it into an ordinary ticket or cross-supplier message.

Do not delay a necessary protective action merely to complete a perfect form. The incident coordinator should record what is available, mark uncertainty and update the record as evidence changes.

Map the service by layer

Use one row per layer or sign-off, even when the same organisation owns several rows. The owners and authority below are examples to replace with the actual supplier, customer and service agreement; the final business-acceptance row is a sign-off, not another technical layer.

LayerExample failureInvestigation ownerAuthority to changeEvidence for handover
Domain registrationDomain suspended or expiredDomain owner or registrar administratorAuthorised registrant/contactRegistrar status and renewal record
DNSWrong or missing answerDNS operatorNamed DNS change approverAuthoritative query, zone change and propagation check
Edge and certificateTLS or proxy failureEdge/certificate ownerProvider or authorised operatorCertificate chain, edge status and origin test
Network and virtual machineHost unreachable or resource unavailableInfrastructure providerProvider operatorPlatform status, reachability and resource evidence
Operating system and runtimeWeb service or database process failedServer/runtime ownerManaged provider or customer administratorService state, logs, configuration and change history
ApplicationError after deployment or plugin updateAgency or application ownerApplication release ownerApplication logs, reproduction and rollback result
External dependencyPayment, identity or API failureIntegration ownerSupplier plus application ownerSupplier status, request identifier and fallback result
Business acceptanceSite loads but checkout remains wrongClient/agency business ownerNamed acceptance authorityFunctional checks and acceptance record

This is a starting structure, not a universal allocation. The NCSC describes cloud responsibility as dependent on the service and its implementation. A managed provider may own some operating-system or runtime work, but that does not automatically move application diagnosis or client communication.

Name five incident roles

Small teams can combine roles, but they should still name the functions.

Incident coordinator

Keeps the shared timeline, confirms priority, assigns actions, resolves ownership gaps and calls decision-makers. This role does not need to perform every technical task.

Platform investigator

Checks the hosting, operating system and named runtime components within the applicable service scope. Supplies evidence when the platform looks healthy or when application input is required.

Application investigator

Checks deployments, code, plugins, queries, configuration and integrations. Confirms whether the business function works after platform recovery.

Decision owner

Can approve consequential actions such as disabling checkout, rolling back a release, restoring older data, changing DNS or accepting a degraded workaround.

Client communication owner

Provides clear updates to the end client. A platform status message is an input to this role, not a substitute for explaining the client's actual service impact.

For disruptive cyber incidents, the NCSC's preparation guidance says operational response, technical recovery, legal/regulatory work, decision-making and communications should have defined owners. It also recommends delegated authority so urgent decisions can still be made if senior leaders are unavailable. Apply the same ownership discipline proportionately to other website incidents; the guidance does not define a Novasean support promise.

Use a fixed investigation sequence

1. Protect people, data and evidence

If the symptom suggests compromise, unsafe transactions or data corruption, use the applicable security plan. Restrict harmful activity where authorised and preserve the records needed for investigation. Do not turn a security incident into an ordinary availability ticket.

2. Establish business impact

Ask what users cannot do. A failed brochure page, unavailable checkout and corrupted client portal require different decisions even if they share the same technical cause.

3. Check broad dependencies before narrow ones

Confirm domain, authoritative DNS, certificate/edge behaviour, provider status and basic reachability. Then move through operating system, runtime, data services, application and external integrations. Parallel checks are useful when they do not risk conflicting changes.

4. Record evidence, not verdicts

“No platform fault found” is weak. “VM reachable; disk at 62%; web service active; origin health endpoint returned 200 at 14:08 UTC; application checkout still returned 500 with request ID…” gives the application owner a usable handover.

5. Control changes

Record who authorised each change, what was changed, the expected result and how to reverse it. Avoid simultaneous uncoordinated changes that destroy the ability to identify what worked.

6. Validate in layers

Technical restoration is not business recovery. Validate infrastructure health, runtime health, application behaviour, data correctness and the affected customer journey. The relevant owner should accept each layer.

7. Close and learn

Capture the cause as far as evidence supports it, the impact, actions, unresolved risks and follow-up owners. Update the responsibility map when the incident reveals a missing dependency or ambiguous handover.

Define the handover rule

Every boundary needs a minimum handover package:

  • incident identifier and current coordinator;
  • observable symptom and business impact;
  • start time and current status;
  • checks completed and their timestamps;
  • the minimum relevant redacted log excerpts, request identifiers or screenshots, shared through an approved restricted route;
  • changes already made;
  • the specific question for the receiving owner;
  • urgency and next coordination time.

Do not put secrets, unnecessary personal data or client-confidential content into an ordinary ticket or cross-supplier message. If an unaltered original must be preserved, keep it separately with a named custodian and access trail; share only what the receiving owner is authorised to see. The receiving owner should acknowledge ownership or explain, with evidence, why the issue needs a joint investigation or another owner. A ticket reassignment without context is not a handover.

Rehearse one ambiguous scenario

Use a hypothetical agency scenario that crosses the boundary; a checkout workload is not a statement about the current Novasean Managed VPS offer:

Checkout returns an error after planned server maintenance. The VPS is reachable, the database process is running and no external payment outage is reported.

Ask each participant what they check, what they may change, what they need from another party and who updates the client. The NCSC recommends exercises that include technical responders and wider organisational roles so dependencies and decisions are tested, not merely discussed.

Record disagreements and fix the map. The purpose is not to prove that every incident will be easy. It is to remove avoidable uncertainty before pressure makes coordination harder.

Where Novasean currently draws the boundary

Novasean's live responsibility page describes the current Managed OS catalogue: scheduled operating-system updates, monitoring set-up, up to one hour of OS administration per month and server backups, with schedule, retention and recovery arrangements agreed during onboarding. Application code, content, changes and functional checks remain with the customer's team unless separately agreed. The page promises neither round-the-clock response nor a recovery target, and the live Managed VPS page says new orders are temporarily paused.

The page also labels the agency portfolio enquiry as a separate, non-binding route, not a VPS plan or accepted service commitment. Neither public page assigns all the incident roles or handover steps in this article. For an actual service, name those people, decision rights, response arrangements and evidence in the applicable agreement and operational plan; this article does not establish them.

Sources

Next: Review the current platform and application boundary.

Keep exploring

Compare the current Managed VPS plans and responsibility boundary, or browse all articles. New VPS orders are temporarily paused.