Sitonce
Country: US
Show exams for United States Hong Kong
Sign in

Azure Architecture Trade-offs: Identity, Data and Resilience

Updated 11 min read
Key takeaway

AZ-305 architecture questions ask you to translate requirements into a design and choose among competing goals such as least privilege, latency, cost, availability, recovery, and operational effort.

  • A reliable method is to separate the requirements, map each to an architecture control, test dependencies, and compare trade-offs.
  • The strongest design meets stated needs without adding unnecessary complexity or leaving a critical requirement uncovered.
On this page12 sections
  1. Architecture begins with requirements, not products
  2. Trade-off 1: identity convenience and least privilege
  3. Trade-off 2: relational fit, scale, and operational model
  4. Trade-off 3: redundancy, backup, and recovery
  5. Trade-off 4: private access and network complexity
  6. Trade-off 5: eventing, messaging, and application architecture
  7. Trade-off 6: observability and governance
  8. Worked architecture exercise
  9. A decision checklist for AZ-305
  10. Common trade-off errors
  11. Summary
  12. Use a decision record to check completeness

Architecture begins with requirements, not products

The current AZ-305 outline measures the ability to design Azure and hybrid solutions. An architect must interpret a stakeholder’s business goal, translate it into technical requirements, and recommend a design. This means the first step is not choosing a service name. It is identifying what the system must do, what it must avoid, and how success will be measured.

Separate functional requirements from nonfunctional requirements. Functional requirements describe behavior: users submit orders, analysts query records, or workers process events. Nonfunctional requirements set qualities and constraints: response time, availability, recovery point, recovery time, privacy, compliance, cost, location, and operational skills. Many plausible designs meet the functional need. The nonfunctional conditions distinguish the best one.

A practical requirements table has four columns: requirement; measurable target; design element; and validation method. ‘Highly available’ is too vague until translated into a failure boundary and an acceptable interruption. ‘Low cost’ needs a budget or comparison. ‘Secure’ needs an identity, network, encryption, and monitoring model. If the case omits a crucial detail, state the assumption rather than silently inventing it.

Trade-off 1: identity convenience and least privilege

Applications and people need identities to access resources. A secret stored in configuration may be familiar but creates rotation and exposure risk. A managed identity can avoid embedding a long-lived credential for supported Azure workloads, but the architecture still needs correct role assignment, scope, and service support. A broader permission may simplify setup while increasing blast radius.

Work out the principal, action, and scope. A web application reading blob data needs a data-plane permission for its workload identity. An administrator changing storage account settings needs management-plane permission. One role may not cover both. Use narrow scope where feasible, then test that the workload can perform the required action and no unrelated action.

Governance adds another layer. Management groups, subscriptions, and resource groups can reflect ownership and policy boundaries. Tags can support organization and reporting. Azure Policy can audit or deny configurations that violate a rule. RBAC controls who can perform management operations. Identity governance manages entitlement lifecycle. These tools complement each other; a tag does not enforce a rule, and Policy does not grant an operator permission.

Worked case: a service principal needs to read one data store, while a platform team can deploy the surrounding resources. A poor design gives the application broad subscription-level Contributor rights. A stronger architecture assigns the application identity the narrow data role at a suitable scope and gives the platform team its own deployment permissions. Secrets and certificates are managed separately. This reduces the effect of a compromised application without blocking legitimate administration.

Trade-off 2: relational fit, scale, and operational model

Data design starts with access patterns and consistency needs. Relational workloads often require structured schemas, transactions, joins, or existing SQL compatibility. Semi-structured or unstructured workloads can benefit from other storage services and APIs. The architect should compare service capabilities, service tier, compute options, scalability, data integration, protection, and cost against the workload rather than choosing a database from popularity.

Managed services can reduce patching and infrastructure operations, but may constrain versions, extensions, control, or migration choices. A higher tier may improve throughput or resilience but incur cost. A distributed data service may improve global access while requiring careful decisions about consistency, partitioning, and conflict behavior. A storage design can optimize for cost yet make retrieval slow or difficult. Identify the access pattern and service-level objectives before selecting a tier.

Worked case: a legacy application depends on a SQL feature that a target managed service does not support. A direct migration to that service may reduce operations but break compatibility. The architect can evaluate an alternate managed database, a temporary VM-based database, or modernization of the application. Compare supportability, downtime, security patching, cost, and the migration timeline. The answer depends on the case constraints, not on a blanket rule that PaaS is always preferable.

Trade-off 3: redundancy, backup, and recovery

Resilience has several parts. High availability keeps a service operating through certain component or zone failures. Replication creates additional copies, but may replicate unwanted changes. Backup provides recovery points with configured retention. Disaster recovery prepares for a broader outage and may require failover, DNS or routing changes, and data reconciliation. These controls are not interchangeable.

Translate recovery point objective into the maximum acceptable data loss and recovery time objective into the acceptable outage duration. Then identify failure scope: host, zone, region, accidental deletion, corruption, or operator mistake. Select protections for each. A requirement to recover deleted data should not be answered solely with zone-redundant storage. A requirement for low downtime should not rely only on a nightly backup.

Higher resilience may increase cost and operational complexity. Cross-region replication can create data residency or consistency issues. Active-active architectures may reduce downtime but require conflict handling and more sophisticated operations. Active-passive designs may cost less while failing over more slowly. The exam scenario tells you which trade-off matters through its target and constraints.

Worked case: a finance system must survive a zone failure with minimal interruption, but can tolerate several hours of regional recovery. Use a zone-resilient primary architecture for the local failure and a separate regional backup or disaster recovery plan sized to the stated recovery objectives. Test restoration and failover. Do not pay for an always-active second region if the case does not need that level of recovery and the added cost is unacceptable.

Trade-off 4: private access and network complexity

Private connectivity improves control over traffic paths, but it introduces DNS, routing, address planning, and operations dependencies. A private endpoint can give a supported service a private IP in a virtual network. Clients need a route to that network and name resolution that directs the service name appropriately. A private endpoint does not automatically connect an on-premises network to Azure or create all required DNS links.

Hybrid connectivity choices involve bandwidth, latency, resilience, routing, provider availability, and cost. A VPN may be suitable for encrypted connectivity with variable performance needs. ExpressRoute may suit private connectivity and more predictable network performance when the business justifies it. The architecture must consider redundancy and failover rather than assuming a single circuit is highly available.

Network security groups filter traffic; firewalls can provide broader inspection or centralized egress controls; identity permissions control resource actions. These operate at different layers. A design that blocks public traffic but gives every workload broad identity rights is not fully secure. A design with least-privilege roles but publicly exposed data may also miss the security requirement.

Trade-off 5: eventing, messaging, and application architecture

Synchronous calls are straightforward when a response is required immediately and downstream services are reliable. They can create tight coupling and propagate a delay or outage. Messaging can buffer bursts and let components process asynchronously, but it adds eventual consistency, retry, duplicate handling, and dead-letter considerations. Event-driven designs can notify multiple consumers but require event contracts and idempotent processing.

Choose a pattern from the workload. If an order API must accept work while a document processor is temporarily busy, a durable queue may help. If the client must receive an immediate eligibility decision, a synchronous call may be required, though timeout and resilience design still matter. If independent services need to react to a published fact, an event architecture can decouple producers and consumers. Do not use a queue merely because it is a modern pattern.

Caching can reduce repeated reads and latency, but it introduces staleness, invalidation, and data protection questions. Application configuration management centralizes settings and supports controlled rollout, while secret storage protects credentials. Automated deployment can improve repeatability but requires environment configuration, identity, approvals, and rollback plans.

Trade-off 6: observability and governance

Logging and monitoring have different but connected jobs. Logs preserve events and support analysis. Metrics expose numerical signals and trends. Traces show request flow across components. Alerts notify or trigger configured action when conditions are met. Security monitoring needs identity and network events, retention, access controls, and investigation workflows. A dashboard alone does not provide data collection or response.

Centralization can support correlation and governance, but collecting everything forever can be expensive and create privacy risk. Decide which sources matter, how they are routed, who can view them, how long they are retained, and what responders do. Compliance requirements may require a separate archive or immutable retention. Operational teams need enough context to distinguish platform issues from application failures.

Worked architecture exercise

A company runs a customer portal on Azure and an order system on-premises during a staged migration. The portal must reach order data privately, support sudden demand peaks, meet a defined recovery objective, and provide central security monitoring. Start with the migration sequence and hybrid network path. Decide how the portal authenticates and what data it can access. Determine whether the order interface is synchronous or can use messaging. Select data and compute patterns based on compatibility and demand. Set backup and failover for the specified recovery target. Route logs and security signals to a governed analysis environment.

Then test the design against changed conditions. If the on-premises order system cannot tolerate a new dependency, use a staged integration boundary. If the recovery objective becomes shorter, reconsider regional architecture and failover automation. If data cannot leave a jurisdiction, confirm storage location and telemetry handling. If the operations team is small, prefer managed capabilities where they satisfy compatibility and control requirements.

A decision checklist for AZ-305

  1. State the business outcome and measurable success criteria.
  2. Separate must-have requirements from preferences.
  3. Map requirements to identity, data, continuity, infrastructure, and monitoring components.
  4. Identify dependencies such as DNS, secrets, routing, service limits, or operational ownership.
  5. Compare at least one plausible alternative on cost, risk, performance, reliability, and complexity.
  6. Validate the design through documentation, prototyping, testing, or migration assessment.
  7. Explain what the design does not guarantee and what must be monitored or tested.

This checklist works for both exam questions and real solution design. It helps prevent attractive service features from distracting you from a requirement. It also makes answers clear: state the recommendation, tie it to the requirement, and note the trade-off.

Common trade-off errors

  • Assuming the highest availability tier automatically meets backup and disaster recovery needs.
  • Choosing the broadest identity role to avoid troubleshooting permissions.
  • Treating a private endpoint as a complete hybrid network architecture.
  • Selecting a distributed database without confirming consistency and access patterns.
  • Using messaging when the business requires an immediate synchronous result.
  • Centralizing logs without deciding on retention, access, or incident response.
  • Adding complexity or cost without a requirement that justifies it.

Summary

AZ-305 design decisions are trade-offs grounded in requirements. Separate identities and scopes, match data services to access patterns, distinguish high availability from recovery, design network paths with their DNS and routing dependencies, and choose application patterns based on coupling and timing. Validate the design and explain its cost and operational consequences. The best answer is the one that satisfies the stated constraints with the fewest unsupported assumptions.

Use a decision record to check completeness

For a larger architecture case, create a brief decision record. State the context and constraints, list the options considered, record the recommendation, and explain the consequences. Include what happens if a component fails, who operates it, how it is monitored, and how recovery is tested. This format catches designs that focus on service selection but omit the people and processes needed to keep the system working.

A design review can also identify assumptions that should become questions. Is data allowed to cross regions? How many transactions occur at peak? Does the recovery point include in-flight operations? Is there an existing identity provider? What does ‘private’ mean for users and service-to-service calls? In an exam item, use the supplied information; in real work, surface unanswered questions before committing.

Consider a customer portal with seasonal demand and strict data residency. A design that places every component in one region may satisfy residency but leave recovery dependent on a manual rebuild. A second region can improve recovery, yet the architect must verify that replicated data remains within the permitted geography and that failover does not create conflicting writes. The sound recommendation states both the resilience improvement and the residency condition that makes it acceptable.

In a timed scenario, rank constraints before choosing a service. A requirement such as “must continue accepting writes during a regional failure” is stronger than a general preference for low cost. If the proposed platform cannot meet that availability behavior, explain the gap and offer a design that can. Avoid adding complexity for requirements the case never states; every extra component adds cost, operational work, and another failure mode.

Common questions

What is the best AZ-305 architecture strategy?

Translate requirements into measurable design goals, compare suitable alternatives, and explain dependencies and trade-offs.

Is replication the same as backup?

No. Replication may improve availability or durability, while backup provides recovery points under configured retention and restore processes.

Does a private endpoint create hybrid connectivity?

No. It provides private access to supported services; the full network path and DNS must also be designed.

What trade-offs should I consider?

Security, reliability, performance, cost, operational complexity, compatibility, and the workload’s specific recovery and access requirements.