<img height="1" width="1" style="display:none;" alt="" src="https://px.ads.linkedin.com/collect/?pid=2826169&amp;fmt=gif">
Start trial

    Start trial

      img-anim-badge-globe-02Replication is configured, and the replica is current. But what happens when the primary stops responding? Does failover happen automatically, or does someone need to promote the standby?

      That gap between having a standby and having PostgreSQL high availability is where most designs sit.

      The consequences can be expensive. In Uptime Intelligence’s 2026 Annual Outage Analysis, 57% of respondents put the cost of their latest major outage above $100,000, while 20% reported losses exceeding $1 million.

      This guide looks at what it takes to keep the database available when something fails, and the architecture you need to make that happen.

      Learn how to achieve PostgreSQL high availability with effective replication, failover strategies, and architecture components to minimize downtime

      What high availability means for PostgreSQL

      High availability is the ability of the database service to keep accepting work when a component fails. It combines redundancy with automatic failover that shifts traffic from the failed component to a healthy one.

      Both halves are required. Redundancy without automated traffic shifting is a spare, not availability.

      High availability vs. disaster recovery

      img-people-discussing-in-front-of-glass-with-sticky-notes-01-variation-01High availability and disaster recovery both help protect database services from disruption, but they solve different problems:

      • High availability
        Keeps the service running through failure by switching to redundant infrastructure.
      • Disaster recovery
        Restores the service after data loss or a larger outage.

      They can use the same mechanisms, but they are not interchangeable.

      A replica protects against node failure but does nothing about a dropped table, because the drop replicates in milliseconds. The same applies to other logical errors: replication can copy the problem straight to the standby. That is why an availability design is not a recovery strategy.

      What PostgreSQL provides and what it does not

      img-person-using-computer-and-learning-icons-01-variation-03PostgreSQL provides many of the core capabilities needed for high availability, but it does not provide a complete high availability solution out of the box.

      PostgreSQL provides:

      • Streaming and synchronous replication
        Keeps a standby server up to date with changes from the primary server.
      • Hot standby
        Allows read-only queries on a standby while it continues receiving changes.
      • Standby promotion
        Allows a standby server to become the new primary when needed.

      But PostgreSQL does not provide:

      • Automatic failover
        Detecting a failure and promoting a healthy standby without manual intervention.
      • Cluster management and leader election
        Coordinating which server should act as the primary.
      • Client routing
        Directing applications and connections to the current primary after failover.

      PostgreSQL high availability therefore combines its built-in capabilities with external components rather than being something you simply switch on.

      Start with the availability target, not the architecture

      Your PostgreSQL high availability architecture should start with how quickly each workload needs to recover and how much data you can afford to lose. These requirements determine the level of availability you actually need.

      Teams routinely reverse this, choosing a topology and discovering afterward what availability it delivers. Instead, define two targets for each workload first:

      • Recovery time objective (RTO)
        How long the database service can be unavailable before the disruption becomes unacceptable.
      • Recovery point objective (RPO)
        How much data you can afford to lose, usually measured in time.

      You can then design the replication, redundancy, and failover approach around those requirements and the service level agreement (SLA) you need to support.

      For regulated workloads, enterprise database requirements, while AI workloads can introduce their own security and availability requirements.

      Annual downtime falls from 3.65 days at 99% availability to 5.26 minutes at 99.999% availability

      Each additional nine costs disproportionately more in infrastructure, complexity, and operational maturity. Moving from 99.9% to 99.99%, for example, cuts the annual downtime allowance from almost nine hours to less than one.

      Planned maintenance counts too. Patching, upgrades, schema migrations, and hardware refreshes use the same downtime budget, making controlled switchover part of the availability design.

      The components of a highly available PostgreSQL architecture

      A highly available PostgreSQL architecture needs redundancy, reliable health checks, a way to choose the primary server, and a routing layer that sends connections to it. Each component protects against a different point of failure.

      Infrastructure alone does not determine resilience. In Uptime Institute’s 2025 Global Data Center Survey, 87% of respondents who had experienced an impactful outage in the previous three years believed changes to how systems were managed, configured, or operated could have prevented their outage.

      87% of data center leaders say better system management, configuration, or operations could have prevented an outage

      Redundancy and standby nodes

      img-arrows-on-board-01-variation-01Redundancy means maintaining at least one standby server with a current copy of the primary server’s data. Streaming replication can keep that standby current, but replication is only the data movement layer — it does not provide failover by itself.

      Where those servers run matters too. Two standbys in the same rack share a power supply and a switch. Two nodes in the same availability zone share a facility. Spread nodes across the failure domains your availability target needs to withstand, such as power, network, storage, facility, or availability zone.

      For multi-site architectures, data residency requirements may also affect where redundant copies can run, while your wider PostgreSQL security practices need to protect that data across each environment.

      Enterprise security controls, including data encryption, also need to extend across each copy of the database.

      Health checking and failure detection

      img-flowchart-diagram-01-variation-02Health checks determine whether the primary server is healthy enough to keep serving queries and when failover should begin. Getting that decision wrong can cause an outage rather than prevent one.

      Failure detection needs to account for:

      • False positives: A slow but healthy server is incorrectly declared unavailable.
      • Flapping: A struggling node repeatedly moves between healthy and unhealthy states, triggering unnecessary disruption.
      • Network partitions: Different parts of the cluster can no longer communicate and may disagree about which nodes are healthy.

      There is a trade-off. Aggressive health checks can speed up automatic failover but increase false positives, while conservative thresholds reduce false alarms but increase recovery time. Monitoring PostgreSQL activity helps teams set thresholds based on how the database actually behaves rather than guesswork.

      Consensus and leader election

      Consensus and leader election determine which standby server should become the new primary after a failure. The cluster needs to make that decision once and agree on the result. Otherwise, two servers can both believe they are primary.

      Connection routing and load balancing

      Connection routing makes sure applications can find the current primary server after failover. A proxy, virtual IP, or DNS can route connections to the right server, while load balancing can distribute read traffic across read replicas where the architecture supports it.

      The routing method also affects recovery time. For example, DNS-based routing can be slowed by clients that cache old records, even when a short TTL is configured.

      Failover, and the ways it goes wrong

      img-pawn-on-board-03-variation-03PostgreSQL failover moves database operations from a failed primary server to a healthy standby server. Successful failover requires more than promoting the standby. Connections drop, uncommitted transactions are lost, applications need retry logic, and the cluster must prevent two servers from accepting writes at once.

      Automatic failover vs. controlled switchover

      Automatic failover and controlled switchover both move operations to another server, but they happen for different reasons:

      • Automatic failover
        Responds to an unexpected failure without waiting for human intervention.
      • Controlled switchover
        Deliberately moves operations between servers, usually for maintenance, so work can be drained and the change managed.

      Promotion is usually the easy part. The harder part is getting every client to follow within the required recovery time, without applications continuing to hold connections to a server that is no longer primary.

      Split-brain, quorum, and fencing

      Split-brain is a failure where two nodes both believe they are primary and accept writes. It can happen when a network partition isolates the primary without killing it, and the standby is promoted on the other side.

      The result is divergent writes that must be reconciled manually, potentially making split-brain worse than an outage. Dedicated audit logging can help establish which writes occurred on each server after an incident.

      Two mechanisms help prevent split-brain:

      • Quorum
        Requires sufficient voting members to agree before promotion.
      • Fencing
        Isolates the old primary from the network or storage so it cannot continue accepting writes.
      Split-brain occurs when two database nodes act as primary; quorum and fencing help ensure only one can accept writes

      Designs with automated failover and no fencing are common and are a bet that partitions will not happen.

      Failback and rejoining the cluster

      Failback is the process of returning a recovered server to the cluster after failover.

      The old primary may have writes the new one never received, so it usually cannot simply rejoin. It needs rebuilding from the current primary, or reconciling with pg_rewind. Decide the failback process in advance.

      Synchronous commit and the durability trade-off

      Synchronous and asynchronous replication balance data protection against performance differently. The right choice depends on your RPO, workload, and the distance between your primary and standby servers.

      • Synchronous replication
        A transaction is not acknowledged until a standby confirms it received the write-ahead log (WAL). This protects committed transactions during failover but adds write latency.
      • Asynchronous replication
        A transaction is acknowledged by the primary server without waiting for the standby. Writes are faster, but recent transactions may be lost if the primary fails before they replicate.
      Synchronous replication protects data with higher latency, while asynchronous replication is faster but risks recent data loss

      Geography makes that trade-off more important. A nearby standby keeps synchronous replication latency lower, but it may share the same failure domain as the primary server. Moving a standby to another availability zone or region improves protection against wider failures, but the greater distance adds latency to every synchronous commit. A common approach is therefore to use synchronous replication with a nearby standby and asynchronous replication with a more distant one.

      The choice should follow the recovery point objective (RPO) for each workload rather than a single cluster-wide rule. Different workloads may warrant different durability settings, and PostgreSQL allows synchronous_commit to be set per transaction.

      Building high availability you can put an SLA behind

      img-people-at-office-18-variation-03Everything above is achievable on community PostgreSQL. The difference is what stands behind your availability commitment when you have an SLA to meet.

      An internal SLA is only as credible as the operational process behind it, particularly for banking and financial services that depend on database availability for critical applications. That means regularly testing failover and recovery, and maintaining the platform so it is ready when a failure happens.

      Fujitsu Enterprise Postgres supports that operational model with:

      • High availability capabilities and 24/7 global enterprise support with defined SLAs
        Giving teams expert support when availability issues arise.
      • Lifecycle support of up to 10 years
        Reducing pressure to upgrade simply to remain supported.
      • Flexible deployment
        Supporting on-premises, cloud, and hybrid environments, with Kubernetes and Red Hat OpenShift options.
      • 100% PostgreSQL compatibility
        Allowing teams to retain familiar PostgreSQL tooling and operational knowledge.

      Failover behavior is difficult to judge from documentation, since what matters is how your applications respond when the primary goes away. Try Fujitsu Enterprise Postgres and rehearse a failover against your own workload.

      Frequently asked questions about PostgreSQL high availability

      What is high availability in PostgreSQL?roundel-circular-arrow-and-arrow-01

      High availability is the ability of the database service to keep accepting work through component failure, using redundant nodes plus automated traffic shifting to a healthy one. PostgreSQL provides the primitives, including replication and standby promotion, but not a complete solution, so the architecture is assembled.

      What is the difference between high availability and disaster recovery?roundel-circular-arrow-02

      High availability keeps the service running through failure. Disaster recovery restores it after loss. They overlap in mechanism and differ in purpose. A replica protects against node failure but replicates a dropped table within milliseconds, so it is no substitute for backups and point-in-time recovery.

      Does PostgreSQL have built-in automatic failover?roundel-gear-with-check-mark-01

      No. PostgreSQL includes replication and the ability to promote a standby, but no failure detection, leader election, or client routing. Automatic failover requires external cluster management components, which is why two PostgreSQL deployments with identical replication can have very different availability characteristics.

      What is split-brain and how do you prevent it?roundel-blueprint-and-pencil-and-esquadro-02

      Split-brain is two nodes both believing they are primary and accepting writes, usually caused by a network partition that isolates the primary without stopping it. The result is divergent data reconciled manually. Prevention combines quorum, requiring majority agreement before promotion, with fencing that isolates the old primary.

      Should you use synchronous or asynchronous commit for high availability?roundel-gear-and-computer-circuitry-03

      It depends on the workload's RPO and the distance between nodes. Synchronous commit prevents losing committed transactions on failover and adds latency proportional to distance. Asynchronous is faster and accepts a loss window. Set it per workload, and consider synchronous to a nearby standby with asynchronous to a distant one.

      Is replication enough to achieve PostgreSQL high availability?roundel-interconnected-dots-01-1

      No. Replication moves data to a standby, which is necessary but not sufficient. Without automated failure detection, a mechanism for deciding which node is promoted, and a routing layer that redirects clients, recovery needs someone to notice and intervene. Replication alone does not provide automatic failover.

      How often should PostgreSQL failover be tested?roundel-clipboard-with-question-mark-01

      PostgreSQL failover should be tested regularly based on the workload and its availability requirements, as well as after any change to the cluster, the routing layer, or the application's connection handling. Test controlled switchover and simulated failure separately, since they exercise different paths. Untested failover tends to fail on client behavior rather than promotion.

      Topics: PostgreSQL

      Receive our blog

      Search by topic

      see all >
      photo-fujitsu-in-hlight-circle-orange-to-yellow-02
      Fujitsu
      We make the world more sustainable by building trust in society through innovation.

      Fujitsu provides migration, support and training services for PostgreSQL, plus Fujitsu Enterprise Postgres, the open source based database with enhanced enterprise capabilities.
      roundel-owl-and-book-01PostgreSQL Insider 
      has a series of technical articles for PostgreSQL enthusiasts of all stripes, with tips and how-to's.
      Explore PostgreSQL Insider >
      Subscribe to be notified of future blog posts
      If you would like to be notified of my next blog posts and other PostgreSQL-related articles, fill the form here.

      Read our latest blogs

      Read our most recent articles regarding all aspects of PostgreSQL and Fujitsu Enterprise Postgres.

      Receive our blog

      Fill the form to receive notifications of future posts

      Search by topic

      see all >