Capacity, reliability, and operational requirements establish the demand a product must support, the outcomes it must preserve, and the work required to keep it operating. This paper explains how Structured Discussion turns Product Intent, Stable Use Cases, organizational obligations, operational evidence, and Site Reliability Engineering (SRE) principles into requirements suited to the product. Participants decide which failures matter, when intervention is justified, how much interruption is acceptable, and which recurring work engineering should eliminate. They preserve those judgments with their rationale, applicability, and decision authority in the Knowledge System.
These Standing Requirements guide architecture, functional specifications, and technical specifications, which translate the approved judgments into implementation and verification. Product Owners or Product Managers approve business commitments and acceptable consequences; architects, Development Leads, or the accountable engineering team establish the supporting technical requirements and designs. SREs and other specialists contribute analysis and professional advice. Each discussion must examine whether its decisions satisfy the requirements it consumes.
The required specificity is reached when a qualified participant can assess a proposal against the recorded obligation, identify the evidence needed for acceptance, and recognize decisions still open to engineering. An unfamiliar-participant test examines whether the Knowledge System supports that independent judgment. This paper sets out the preparation activities, decision records, environment-specific observation requirements, and validation needed to carry product judgments through to code implemented by engineers or AI coding agents.
1. Illustrative Product Context
One online shopping platform is used throughout the paper. Customers browse products, maintain carts, and place orders; the marketing team runs campaigns and uses a daily report populated from vendor sales data to allocate spending. The capacity example assumes 10,000 purchasing customers per hour, each placing one order averaging $50. Applicable organizational policy requires cloud hosting. The product team must decide which demand to support, when a data discrepancy prevents use of the report, and which interruptions to purchasing are acceptable.
The same product connects the later examples. Automated vendor-data checks can reduce repeated investigations by applying known campaign conditions. A checkout resilience requirement can lead to an AWS-and-Azure architectural direction, supported by functional behavior during interruption and technical details for deployment, routing, state, recovery, and instrumentation. The figures and design choices illustrate decisions for this product; the framework responsibility is to establish requirements that later participants can apply and verify.
2. Product-Specific Requirements Through Structured Discussion
2.1. Inputs, Interpretation, and Decision Authority
Requirement formation begins with the product situation. Product Intent explains the purpose being served; Stable Use Cases identify recurring user activities and outcomes. Applicable policy supplies obligations such as permitted hosting environments or restrictions on data replication. Existing requirements, architecture knowledge, incident records, workload forecasts, and operational practices reveal commitments and assumptions already affecting the product. SRE principles contribute questions about acceptable risk, service objectives, monitoring, capacity, and repetitive operational work.
The participants use Structured Discussion to interpret these inputs together. For each concern, they identify the affected use case, the relevant operating conditions, the consequence of failure, the required response, and the person authorized to decide. Conflicting inputs become explicit decisions. A campaign forecast may increase required capacity while a cost constraint limits redundancy; a data-handling policy may restrict the recovery location that an architect proposes. The participants must resolve how both statements apply to the product and record the authorized trade-off.
Eliciting architectural quality concerns before selecting a design gives participants concrete conditions against which to evaluate their options. A team may use interviews, scenario analysis, incident review, or a facilitated workshop to establish those concerns. Quality Attribute Workshops (QAWs) are one documented method: stakeholders examine business and technical context, generate and prioritize scenarios, and refine selected scenarios through the stimulus source, stimulus, environment, affected artifact, response, and response measure.1 A scenario can therefore express which failure occurs, under what conditions, and what measurable response is required. Specification-First Delivery accepts any elicitation method that produces sufficiently explicit, approved judgments for subsequent work. QAW is an optional reference for scenario refinement. The resulting product commitment must also identify its decision owner and the authoritative record later contributors will use.
Capacity and reliability each involve business and technical decisions. The following comparison identifies the authority and output of those decisions.
Business capacity
The Product Owner or Product Manager decides the demand the product must support, including customer activity, order volume, campaign peaks, planning horizon, and permitted demand restrictions. Forecasts inform this commitment and retain their assumptions and uncertainty.
SREs, security engineers, data specialists, and quality professionals contribute failure analysis, feasibility evidence, measurement design, and operating experience. Their participation improves the decision; business commitments remain with the Product Owner or Product Manager. Existing security and policy approval rights still apply. Engineering may decide technical requirements by team consensus when the team identifies who will record and maintain the result. If a feasible design would change an accepted business consequence, the product decision must be reopened with its owner.
2.2. Specificity Sufficient for the Next Decision
A requirement is sufficiently specific when its conditions, required outcome, evaluation criteria, and consequential response support the next dependent decision. The record defines the required result precisely enough to assess a candidate mechanism, while the appropriate technical authority chooses that mechanism.
For the shopping platform, a requirement that campaign reporting must withhold an unexplained vendor-data discrepancy before budget decisions are made identifies an affected activity and required behavior. Completing it requires the comparison population, intervention criterion, decision deadline, permitted fallback, and release authority. Those details determine what a design must accomplish. The choice of telemetry backend, notification API, or validation library can then be made in architecture or technical specification work.
The preparation activity produces an approved record containing the applicable use case and conditions, required outcome, measurement or review method, response to a deviation, decision owner, rationale, and review triggers. Where a numerical target matters, the record includes its population and evaluation window. Where professional judgment remains necessary, it identifies the evidence and authority for that judgment. Persistent obligations enter the product's Standing Requirements.
Requirement translation
From organizational obligations to code
- 1
Policy or organizational requirement
Establish the source obligation, its authority, and the conditions under which it applies.
- 2
Capacity, reliability, and operational requirements
Use Structured Discussion to interpret the obligation alongside product intent, use cases, evidence, and SRE principles, then approve the product-specific commitment.
- 3
Architecture specification
Select system responsibilities, dependencies, and design constraints that satisfy the approved requirements.
- 4
Functional specification
Define observable behavior, including degraded operation, exceptions, recovery outcomes, and intervention.
- 5
Technical specification
Translate the approved architecture and behavior into implementation work, configuration, instrumentation, and verification with explicit scope and authority.
- 6
Code
Engineers or AI coding agents implement the specified behavior and controls, with evidence linking the result to the originating requirements.
This sequence shows where the Standing Requirements go: an organizational obligation becomes a product judgment, then an architectural arrangement, specified behavior, executable technical work, and code. The specifications retain the applicable requirements and decision provenance as they add detail. Teams may refine related specifications together; the sequence expresses the translation of an obligation through the delivery artifacts.
For work assigned to an AI coding agent, the technical specification and its linked context must make every applicable requirement discoverable. A requirement is carried through completely when each material obligation has an implementation responsibility and a way to verify the resulting behavior. An agent may implement the entire authorized code change, while accountable participants review conformance and acceptance evidence. Traceability and verification establish whether the obligation reached the implementation.
The unfamiliar-participant test can expose omissions while requirements are still being formed. A participant who cannot determine whether a campaign explains an alert has identified missing product context or an unresolved judgment. The owner supplies or decides that information, and the affected part of the test is repeated. The test contributes evidence of usability; approval comes from the responsible human authority.
3. Adapting SRE Principles to the Product
3.1. Reducing Toil as a Requirement
Operational requirements must account for the human work a design creates. In Google's SRE treatment, toil is production work characterized by repetition, manual effort, potential automation, limited lasting value, and growth with service scale. Engineering work that removes a recurring operational task creates a lasting improvement. This account makes repetitive intervention an object of system design, with consequences for the engineering time available to improve reliability.2
The shopping platform's team applies that principle by examining the recurring tasks its operation would require. Its vendor-data pipeline may involve checking daily totals, finding the campaign calendar, dismissing expected alerts, and rerunning the same comparison. Its checkout operation may involve manually verifying every deployment in two clouds or repeatedly preparing a standby environment. Structured Discussion asks which decisions require human judgment, which checks can be automated, and which missing context causes the same investigation to recur.
The resulting operational requirement should state the expected behavior. For its reporting pipeline, the team could require routine imports to be validated automatically against the approved comparison rules, recognized campaign conditions to be applied without repeated manual triage, and unresolved discrepancies to reach an accountable reviewer with the evidence needed to decide. That requirement changes the design: the validation process needs access to maintained campaign context and must retain its decision evidence. The record names the recurring checks to automate and the evidence to retain, making the toil-reduction principle usable in design.
Engineering records the expected manual effort, its cause, and how it changes with workload. The Product Owner or Product Manager approves any trade-off affecting staffing, response coverage, cost, or product commitments. A deliberate manual exception can remain appropriate when judgment is necessary and its cost is accepted. Repeatedly performing a predictable check is a candidate for engineering improvement. Recording both cases allows the team to revisit the accepted operating effort when demand grows.
3.2. Service Objectives, Risk, and Response
Reliability targets express a product trade-off. For this platform, the discussion begins with the effect of interrupted purchasing, unavailable browsing, delayed order confirmation, or stale inventory. Different activities may justify different requirements. The SRE treatment of risk relates the chosen level of reliability to user expectations, business consequences, engineering cost, and the opportunity to develop features, with product owners participating in the target decision.3
A service-level indicator (SLI) measures a selected aspect of service behavior; a service-level objective (SLO) sets a target for that indicator. Selecting the measure and evaluation window gives an availability statement its meaning. The SRE guidance emphasizes indicators that represent the service experienced by its users.4 A checkout success target, for example, needs an eligible request population, a definition of success, the treatment of dependency failures, and a measurement location. The measure must establish whether the customer can complete a purchase, including the dependencies that determine that outcome.
An error budget expresses the unreliability permitted by the SLO. Its operational value depends on an agreed policy for action when the budget is being consumed or exhausted. The SLO implementation guidance connects objectives with stakeholder agreement, technical feasibility, ownership, and continuing review, and describes an error-budget policy that directs decisions.5
For the shopping platform, a proposed policy might suspend discretionary checkout deployments after exhaustion of the agreed error budget, permit changes required for recovery, and require product and engineering review before ordinary deployment resumes. The discussion must resolve who classifies an exception and what evidence permits resumption. Once approved, those judgments belong in the Standing Requirement. Technical specifications implement the measurement and release controls needed to apply them.
3.3. Actionable Observation and Alerting
Observation should support a defined operational decision. The product team identifies which user consequence warrants interruption, who can act, and what evidence they need. Monitoring then provides both externally observed service behavior and internal measurements for diagnosis. Google's monitoring guidance develops that relationship and emphasizes alerts that are understandable, actionable, and low in noise.6 For checkout, an alert should direct attention to the threatened purchasing outcome and give the responder enough context to investigate or recover it.
Alert design also has measurable trade-offs. The SRE Workbook evaluates SLO alerting through precision, recall, detection time, and reset time, and develops error-budget burn-rate approaches.7 Those methods apply to an SLO with defined good and bad events. A day-over-day data-volume change must first be related to defined data-quality failures before an SLO burn-rate alert can be derived from it. The data-quality example below starts with fitness for a business decision and derives its intervention rules from that use.
3.4. Telemetry and Log Emission by Environment
Telemetry emission is the production of signals such as metrics, traces, and events by application or infrastructure components. Log emission records discrete events with enough context to reconstruct what occurred. These signals make required behavior observable. A dashboard or alert can use only the evidence the system emits and the observation path retains; requirements must therefore state the expected coverage and level of detail in development and production.
The technical decision makers, supported by SRE and engineering specialists, establish which activities emit signals, their frequency or event coverage, required fields, sampling rules, and permitted diagnostic detail. They also define how quickly the signals must be available, how long they remain usable, who can access them, and what happens when collection fails. Product ownership confirms requirements affecting business intervention or accepted operating effort, and applicable security and privacy policy governs emitted content. Quantitative limits must support the product's detection, investigation, and capacity needs.
For the shopping platform, the following illustrates an approved environment-specific emission contract. The same event meanings and correlation fields apply in both environments; the permitted diagnostic detail and collection settings differ.
Development emission
Make each selected test scenario diagnosable and verify the signals required in production.
- Metrics and traces
- Emit checkout attempt, result, duration, retry, and dependency-failure signals in automated scenario tests. Capture every trace in the selected diagnostic test runs, including calls across service and cloud interfaces. Identify the environment, service revision, and correlation identifier.
- Log detail
- Emit structured INFO records for import completion, validation decisions, and recovery transitions; WARN for recoverable anomalies; and ERROR for failed operations. Allow DEBUG detail for a selected component or test run, using synthetic or appropriately protected data and the approved redaction rules.
- Required evidence
- Tests inspect emitted fields, levels, and correlation, reproduce a failed checkout and a vendor-data exception, and verify that a responder can follow the selected scenario. Development retention and collection limits must cover the agreed investigation period.
Production emission
Support service measurement, business intervention, and incident diagnosis at the supported workload.
- Metrics and traces
- Account for every eligible checkout attempt and outcome in the agreed SLI measurement. Emit latency, traffic, error, and saturation measurements at the approved resolution. Retain a diagnostic trace for each failed checkout; sample successful checkout traces at the approved rate. Record the environment, cloud, service revision, and correlation identifier.
- Log detail
- Emit structured INFO records for import completion, validation decisions, recovery transitions, and release decisions; WARN for recoverable anomalies; and ERROR for failed operations. Keep routine DEBUG emission disabled. Any temporary increase has an authorized scope, duration, and rollback setting. Required operational decision records remain complete.
- Coverage and operating limits
- Each import records its comparison population, counts, completeness result, applicable campaign rule, and exception. Define numeric freshness, retention, sampling, and volume limits before dependent implementation. Signal collection must support diagnosis from the surviving cloud during the covered provider outage.
An emission contract must resolve settings that affect feasibility before implementation depends on them. For example, a coarse measurement interval may hide an interruption material to the approved availability objective. Sampling successful diagnostic traces can manage volume while preserving complete SLI measurement; sampling the only source of required checkout outcomes would leave that objective unsupported. Logging levels describe diagnostic detail, while the record identifies the events whose evidence must be retained.
Architecture discussion determines how required signals remain available under the approved failure scenarios. Functional specifications state observable events, such as a held report or a recovery transition. Technical specifications define instrumentation points, field schemas, exporters, severity mapping, sampling configuration, retention, and tests of the emission contract. Those tests also examine exporter failure, redaction, and overhead under load. A collector's failure must follow the approved product behavior, including any obligation to preserve a release record, and its effect must be observable.
4. Vendor-Data Quality and Product Intervention
4.1. From an Inherited Threshold to an Approved Rule
Suppose the shopping platform's vendor-data pipeline already raises an alert when today's loaded sales-record total differs from yesterday's by 10 percent. The developer selected this rule before the product team established the reporting requirement. The marketing team's campaigns can explain substantial changes, while a smaller unexplained discrepancy may affect its spending decisions. Structured Discussion must establish which conditions require intervention before the data are used in the daily report.
Data quality depends on the use made of the data. Wang and Strong's study elicited quality attributes from data consumers and developed a framework that includes contextual quality alongside intrinsic, representational, and accessibility concerns. Its contextual dimension associates quality with the task for which data are used.8 For this pipeline, the discussion must therefore establish the consumer's decision and deadline before selecting a threshold.
The Product Owner or Product Manager, supported by the relevant business and data specialists, resolves the following questions:
Which business decision uses the data, and what is the consequence of using an incorrect or incomplete total?
Which population, time zone, business date, and completeness conditions make two totals comparable?
Which campaign or vendor conditions explain an expected change, and who maintains that context?
Which unexplained deviation requires review, delayed publication, a qualified report, or another approved response?
Who can authorize use of the data, and what happens if the issue remains unresolved at the decision deadline?
The platform's product owner then approves the following illustrative reporting contract. Its thresholds, times, and response show the required specificity for this product.
Release of the daily campaign-spending report
Marketing uses the report at 09:00 UTC to allocate campaign spending. Only data satisfying the approved validation rules may enter the report automatically.
- Population and completeness
- Count accepted vendor sales records for the preceding UTC business day. Require the vendor completion marker by 08:00 UTC. A missing marker creates a completeness exception and prevents automatic release.
- Volume comparison
- Compare the completed daily count with the preceding complete UTC day's count. An absolute change greater than 10 percent requires review. A missing or zero baseline creates a comparison exception.
- Campaign context
- A product-approved campaign record may replace the day-over-day volume expectation with a specified range for named segments and dates. Completeness and integrity checks still apply. Missing or expired campaign context provides no override.
- Intervention and authority
- Hold a report with an unresolved exception and notify the data operations owner and campaign decision owner. Release requires recorded approval from the Product Owner or an explicitly delegated business decision maker. At 09:00 UTC, an unresolved report remains withheld and its status is communicated to marketing.
- Operational requirement
- Evaluate routine comparisons and applicable campaign expectations automatically. Give reviewers the counts, completeness status, applicable rule, and campaign evidence. Record manual intervention and recurring causes for review.
The approved 10 percent rule now has a purpose, comparison, response, and owner. If the platform's spending decisions were sensitive to a 1 percent discrepancy, the product owner would need to approve that criterion and engineering would evaluate its implications. Neither value is a Specification-First Delivery condition. A forecast-based range may also be appropriate when seasonal behavior makes yesterday an inadequate baseline. Research on automated data-quality verification demonstrates how validation constraints can be combined with historical quality metrics and anomaly detection, including time-dependent expectations.9 Selecting the applicable method and interpreting an exception remain product and technical decisions.
4.2. Architectural Consequences and Operating Effort
Resolving intervention conditions early enables engineering to evaluate the required observation and response. If the platform needs investigation of small, persistent changes, its reporting pipeline may require historical telemetry, segmentation, and visualization. If the approved response is a notification after a completed batch breaches a stable rule, an automated validation job and notification API may be adequate. Response urgency, data availability, expected variation, and diagnostic needs determine the arrangement together with the threshold.
Engineering chooses the observation and intervention method from the required response, evaluation frequency, and diagnostic evidence. Predictable comparisons and responses can be automated even at a 1 percent threshold. Human investigation remains available for cases requiring interpretation. Grafana or an OpenTelemetry-based pipeline may implement parts of the observation design, but the product requirement is the evidence and intervention needed for the use case.
The same discussion should expose the cost of false alarms. If every marketing event requires an engineer to rediscover an approved explanation, the operating process is missing reusable context. Recording campaign applicability and applying it during validation is one possible engineering response. The product team must also decide who owns that context and how expired or missing campaign information is handled. Reducing toil thus influences the requirements, architecture, and maintenance responsibilities together.
5. Capacity and Reliability for an Online Shopping Platform
5.1. From Business Demand to Technical Capacity
Business capacity specifies the activity that the platform must support: browsing, searching, maintaining carts, and purchasing under ordinary and campaign demand. The Product Owner or Product Manager approves the workload commitment, its planning horizon, and any permitted restriction. A forecast supports that commitment while preserving uncertainty. Google's SRE account of capacity planning combines demand forecasts with provisioning needs and load testing that relates infrastructure capacity to service capacity.10
The technical team translates the approved workload into resource demand. E-commerce workload characterization has examined customer sessions and transitions among activities to model how users exercise a site.11 That relationship matters because 10,000 purchasing customers per hour leaves browsing, search, abandoned carts, and burst behavior unspecified. These activities also consume resources. Engineering must obtain or state the assumptions needed to derive request rates, concurrency, storage, and dependency demand.
For example, database growth can be estimated from orders per day, stored bytes per order, retention, indexes, replication, and associated events. Each factor should have a source or a provisional assumption and an owner. If resilience requires one cloud to carry the supported workload while the other is unavailable, capacity estimates must include that condition. A design sized only for the normal division of traffic leaves the failure requirement unsupported.
5.2. Business Exposure and Technical Objectives
With the illustrative workload established above, one hour of interrupted purchasing exposes $500,000 of gross sales: 10,000 purchasing customers multiplied by one $50 order each. This is an estimate of activity affected, not a prediction of realized loss. Customers may defer purchases, abandon them, or complete them through another channel. If the figure describes visitors, a conversion assumption is also required.
The estimate supplies a reason to discuss interruption tolerance. The product decision maker considers when interruption is most consequential, which degraded service remains useful, and what recovery or reconciliation customers require. Engineering then proposes measurable objectives, capacity margins, recovery behavior, and operating effort. Any incompatibility between affordable technical options and the business commitment returns to the appropriate product or policy decision maker.
A proposed target of 99.9999 percent availability needs especially precise interpretation. For a time-based indicator over a 30-day window, that target permits 2.592 seconds of unavailability. A request-based indicator instead limits failed eligible requests and has no automatic conversion to downtime. The team must agree on measurement resolution, covered journeys, dependency treatment, and the evidence required to assess feasibility. The numerical target alone supplies neither a failure model nor a viable recovery design.
6. Recording Judgments in the Appropriate Context
The product team maintains the relationship between an approved obligation and the decisions made to satisfy it. Standing Requirements state what remains required across applicable increments. An architecture specification records the system arrangement selected to satisfy those obligations, within the maintained architecture knowledge. A functional specification defines the observable behavior, including failure and intervention. A technical specification turns the reviewed decisions into implementation work and verification. Operational procedures make the resulting behavior executable in service.
The following records allocate those responsibilities for the shopping-platform example. They describe related knowledge that may be maintained in an existing repository, architecture system, or governed document platform.
Standing reliability and operational requirements
Record the approved product obligation and the conditions under which it applies.
- Example judgment
- Cloud hosting is required; checkout must tolerate the specified provider-outage scenario; the approved availability indicator, target, and window apply; routine recovery work must meet the agreed automation and human-response requirements.
- Authority and supporting context
- Product ownership approves business consequences; technical authority establishes supporting technical requirements. Retain applicable policy, workload assumptions, measurement rules, rationale, and review triggers.
Architecture specification and architecture knowledge
Record the selected system arrangement, its rationale, and its implementation status.
- Example judgment
- Approve containerized application workloads on AWS and Azure, with explicit decisions for state, routing, identity, dependencies, failure capacity, and recovery. Identify any existing single-provider dependency that prevents conformance.
- Authority and supporting context
- The architect, Development Lead, or accountable team records alternatives, constraints, unresolved feasibility questions, and the requirements each decision serves. Approved Direction remains separate from Current architecture until implementation is established.
Functional specification
Define the customer and operator behavior that satisfies the approved obligation.
- Example judgment
- Define checkout behavior during interruption, prevention of duplicate accepted orders, customer-visible recovery status, and the withholding and authorized release of a vendor-data report.
- Authority and supporting context
- Product and domain decision makers confirm the behavior and exceptions against the Standing Requirements and architecture specification; the record identifies their acceptance conditions.
Technical specification
Define the authorized implementation and verification for a delivery increment.
- Example judgment
- Specify deployment configuration, cross-cloud traffic management, health evaluation, network connectivity, replication behavior, environment-specific telemetry and log emission, rollout and rollback, and the tests required for the affected change.
- Authority and supporting context
- The specification identifies the applicable requirement and architecture revisions, permitted engineering decisions, acceptance conditions, and the people responsible for approval and evidence review.
Code and verification
Implement the reviewed decisions and produce evidence for each applicable obligation.
- Example judgment
- Engineers or AI coding agents implement recovery behavior, order reconciliation, report release controls, and instrumentation within the technical specification's scope and authority.
- Authority and supporting context
- Link each material requirement to its implementation responsibility and verification evidence. Accountable reviewers assess conformance and unresolved deviations before acceptance.
Operational procedures and evidence
Make the approved response executable and retain what demonstrates its behavior.
- Example judgment
- Define automated recovery, escalation, human intervention, and restoration checks; retain service measurements, failure-test results, and recurring manual-work records.
- Authority and supporting context
- Named operational owners maintain the procedures. Product and technical authorities review evidence affecting commitments, architecture conformance, or the accepted operating effort.
Each artifact should link to the authoritative decisions it consumes. A later specification can reference the approved availability definition instead of creating another version with a different window. When a technical investigation discovers that the target requires a different failure model or operating commitment, it creates an issue for the appropriate requirement or architecture discussion. Updating an implementation detail cannot silently authorize a weaker product obligation.
7. Cross-Cloud Requirements, Architecture, and Technical Specification
7.1. Detecting an Incompatible Architecture
Assume that the shopping platform's requirement discussion has established cloud hosting, resilience to the loss of either of two cloud providers for the covered checkout path, and an approved 99.9999 percent availability objective. Product ownership has approved the business commitment, and technical authority has defined the supporting objective and its measurement. The requirement record identifies the covered outage and the treatment of orders and customer-visible effects during recovery. These decisions establish the required result; conformance must be demonstrated through the agreed evidence.
An architecture proposal then places all checkout state in one Azure Storage account. During Structured Discussion of the architecture, the participants load the applicable reliability requirement and trace the checkout path through its dependencies. If an Azure outage makes the only usable state unavailable, that proposal cannot satisfy the required provider-outage behavior. The conflict belongs in the architecture discussion before implementation is authorized.
The review evaluates whether checkout state remains usable under the covered provider-outage scenario. Replication across Azure zones or regions can address failures at those levels; the requirement to operate during an Azure-provider outage requires an independently usable path in the surviving cloud. A storage account used for a noncritical or independently recoverable purpose may still be compatible with the requirement. The relevant judgment concerns the dependency's effect on the covered customer activity.
The discussion records the rejected assumption, the affected requirement, and the design decision still needed. If the existing system already has that dependency, Current architecture should show it accurately, with the unmet requirement recorded. The target and the current operating capability have different evidentiary status.
7.2. Approving an Architectural Direction
The architect, Development Lead, or accountable team might select containerized application workloads running on both AWS and Azure. In this example, a container-only application deployment rule becomes an approved architectural constraint because it supports the chosen portability and operating model. The scope of that rule must be explicit: application containers still depend on networking, identity, storage, and other platform services.
Deploying the same container image in both clouds demonstrates a form of deployment portability. The architecture must additionally explain how the surviving deployment obtains usable state, accepts traffic, authenticates requests, completes or reconciles orders, and supports the required load. A common identity service, routing control, release defect, or unavailable state store can affect both deployments. Architecture review follows these dependencies to determine whether checkout remains usable during the approved failure scenario.
The architecture decision should establish the traffic model, state and consistency strategy, capacity during provider loss, permitted dependencies, and responsibilities for recovery. It should also explain the cost and toil of operating two environments. Automated configuration checks and recovery may be necessary to meet the operational requirement; duplicating routine manual work in two clouds can undermine it.
A decision supported sufficiently for implementation enters Approved Direction with its rationale, adoption scope, and remaining validation obligations. An option whose feasibility is still being investigated retains its proposal status. Implementation and evidence determine when the current-state record can be updated. Conformance to the availability objective requires the agreed failure tests and service measurements across the approved evaluation window.
7.3. Specifying Functional Behavior
The functional specification applies the approved resilience requirement to purchasing behavior. It defines what customers see during an interruption, how an in-progress order is identified, when the platform confirms acceptance, how a retry relates to that order, and how recovery reconciles uncertain outcomes. The product decision makers confirm those rules against the requirement, with engineering checking compatibility with the selected state and traffic model. Telemetry and logs must identify the corresponding outcomes so verification and operation can determine whether the rules were followed.
7.4. Specifying Deployment and Verification
A technical specification for the approved direction defines the change at implementation level. It must resolve details whose failure would prevent the architecture from satisfying its requirement. For the cross-cloud increment, these include:
Deployment behavior. Identify the application artifacts, cloud-specific configuration, secrets and identity dependencies, rollout sequence, compatibility checks, and rollback conditions.
Traffic and health decisions. Specify the external entry path, load-balancing or failover mechanism, health signals that represent usable checkout, detection behavior, routing convergence, and treatment of active requests.
Connectivity and state. Define network paths, routing and access rules, replication or reconciliation behavior, consistency expectations, and operation during loss of a link or provider.
Instrumentation and emission. Specify metrics, trace and log fields, emission points, severity, sampling, export, and retention settings for development and production, including observation from the surviving cloud.
Acceptance evidence. Establish tests for provider isolation, dependency failure, interrupted replication, surviving-cloud capacity, duplicate or incomplete orders, restoration, required human intervention, and the emitted evidence of each outcome.
For this platform, the connectivity section could specify redundant site-to-site VPN paths between AWS and Azure, their routing and access rules, and the behavior required when a tunnel fails. Azure Virtual Network peering can connect the relevant networks within Azure; cross-cloud traffic needs its own interconnection. Engineering evaluates the selected paths' throughput, convergence, and failure behavior against the approved architecture and records the configuration and tests in the technical specification.
The acceptance plan measures the covered customer outcome during failure and recovery, including data correctness and operating effort. Fault tests can expose a shared dependency or a slow transition and establish behavior under the tested conditions. Evidence for a long-term availability objective also depends on the agreed operational measurement and review process. A successful deployment to both clouds supplies only part of that evidence.
8. Validation Through an Unfamiliar-Participant Test
The unfamiliar-participant test evaluates whether the recorded knowledge supports an independent professional judgment. Knowledge System readiness establishes the wider preparation responsibility. For reliability work, select a qualified person who did not participate in the original decisions and give that person the normal knowledge entry point, the proposed change, and the access expected of a real contributor. Evaluate an AI contributor through its actual retrieval path as well; human access does not establish that an agent can discover the same material.
For the shopping-platform example, present the architecture that deploys application containers in both clouds while retaining a critical dependency on one Azure Storage account. Ask the participant to assess it using the maintained sources. The expected result is a reasoned finding that deployment portability leaves the provider-outage requirement unresolved, with a reference to the applicable obligation and the technical authority responsible for the architecture decision.
The test should examine whether the participant can:
find the applicable business commitment, technical reliability requirement, and approved workload conditions;
distinguish Current architecture, Approved Direction, and unresolved proposals;
identify the dependency that conflicts with the required failure behavior and explain its customer consequence;
identify the functional behavior, technical specification work, and evidence needed to resolve the conflict;
determine what telemetry and logs must be emitted in development and production, and whether the proposed observation path survives the covered failure; and
route an unresolved product, architecture, or policy decision to its owner without inventing approval.
The same platform's reporting operation offers a second assessment. Give the participant a threshold breach during a documented campaign and ask whether the daily spending report can be released. The knowledge should support the answer through comparison rules, campaign applicability, completeness checks, and release authority. If the participant must ask someone from the original meeting what the threshold meant, the result identifies missing or inaccessible context.
Record each finding as a retrieval gap, an ambiguous judgment, a conflicting decision, or missing evidence, with an owner and the dependent work affected. The responsible person updates the authoritative source or resolves the decision, and the participant repeats the relevant assessment. A pass means the participant can reach a justified conclusion and identify unresolved authority correctly. It demonstrates usable knowledge; system conformance still requires the agreed technical and operational evidence.
A material gap blocks the decision that depends on it. An unresolved provider-outage requirement prevents approval of an architecture that claims to meet it. Engineering can continue an explicitly authorized feasibility investigation while the product commitment remains open. The person who discovers the gap can describe its consequence and proposed resolution; decision authority remains with the accountable owner.
9. Maintaining Requirements and Their Consequences
The product team maintains the approved requirement, its rationale, linked architecture decisions, implementation evidence, and operating feedback as related knowledge. Review is triggered by changes that affect those judgments: a new campaign pattern, changed vendor delivery behavior, revised demand, repeated manual intervention, altered dependency behavior, or a failure test that contradicts a design assumption.
For example, growth in unexplained pipeline alerts should lead the owners to examine the comparison rule, campaign context, and manual review burden. Raising the threshold changes a business intervention decision and requires its owner's approval. A failed cross-cloud recovery test instead first raises an architecture or implementation question; if the original objective proves impractical, the business commitment returns to product discussion. The source of the change determines which judgment must be reopened.
Security and privacy requirements remain applicable when resilience introduces additional copies of data, network paths, identities, or diagnostic records. The participants reconcile those constraints with the recovery design and record any authorized exception through the appropriate process. Any proposed operational exception must retain the applicable policy approval and its rationale.
The maintained records should let the next contributor follow a product concern through its Standing Requirement, architecture specification, functional specification, technical specification, code, and evidence of behavior. Structured Discussion establishes and revises those judgments. The unfamiliar-participant test shows whether another contributor can use them. Operational evidence then returns actual failures, workload changes, and recurring toil to the same decision process.
References
- Barbacci, M. R., Ellison, R. J., Lattanze, A. J., Stafford, J. A., Weinstock, C. B., and Wood, W. G. (2003). Quality Attribute Workshops (QAWs), Third Edition. CMU/SEI-2003-TR-016. Software Engineering Institute, Carnegie Mellon University. Report and DOI.
- Rau, V. (2016). “Eliminating Toil.” In B. Beyer, C. Jones, J. Petoff, and N. R. Murphy (Eds.), Site Reliability Engineering: How Google Runs Production Systems. O’Reilly Media. Chapter.
- Alvidrez, M. (2016). “Embracing Risk.” In B. Beyer, C. Jones, J. Petoff, and N. R. Murphy (Eds.), Site Reliability Engineering: How Google Runs Production Systems. O’Reilly Media. Chapter.
- Jones, C., Wilkes, J., and Murphy, N. (2016). “Service Level Objectives.” In B. Beyer, C. Jones, J. Petoff, and N. R. Murphy (Eds.), Site Reliability Engineering: How Google Runs Production Systems. O’Reilly Media. Chapter.
- Thurgood, S., and Ferguson, D. (2018). “Implementing SLOs.” In B. Beyer, N. R. Murphy, D. K. Rensin, K. Kawahara, and S. Thorne (Eds.), The Site Reliability Workbook. O’Reilly Media. Chapter.
- Ewaschuk, R. (2016). “Monitoring Distributed Systems.” In B. Beyer, C. Jones, J. Petoff, and N. R. Murphy (Eds.), Site Reliability Engineering: How Google Runs Production Systems. O’Reilly Media. Chapter.
- Thurgood, S. (2018). “Alerting on SLOs.” In B. Beyer, N. R. Murphy, D. K. Rensin, K. Kawahara, and S. Thorne (Eds.), The Site Reliability Workbook. O’Reilly Media. Chapter.
- Wang, R. Y., and Strong, D. M. (1996). Beyond accuracy: What data quality means to data consumers. Journal of Management Information Systems, 12(4), 5–33. Full text.
- Schelter, S., Lange, D., Schmidt, P., Celikel, M., Biessmann, F., and Grafberger, A. (2018). Automating large-scale data quality verification. Proceedings of the VLDB Endowment, 11(12), 1781–1794. DOI: 10.14778/3229863.3229867.
- Benjamin Treynor Sloss. (2016). “Introduction,” subsection “Demand Forecasting and Capacity Planning.” In B. Beyer, C. Jones, J. Petoff, and N. R. Murphy (Eds.), Site Reliability Engineering: How Google Runs Production Systems. O’Reilly Media. Chapter.
- Menascé, D. A., Almeida, V. A. F., Fonseca, R., and Mendes, M. A. (1999). A methodology for workload characterization of e-commerce sites. Proceedings of the First ACM Conference on Electronic Commerce, 119–128. DOI: 10.1145/336992.337024.