Stabilizing a Compromised WordPress Platform During an Active HTTP Flood

Incident Response Under Active Attack

An engineering case study in evidence preservation, controlled experiments, and temporary request-path isolation on a compromised e-commerce origin under an active HTTP flood.

Executive Summary

The initial symptom looked like an application compromise: a production e-commerce platform was selectively serving or redirecting visitors toward malicious casino content.

During the investigation, a second problem emerged.

A distributed HTTP flood was concentrating almost entirely on the homepage. Measurements during high-intensity intervals were approximately 2,000–2,500 requests per second, reaching an origin with one vCPU and approximately 2 GB of RAM.

Recovering WordPress was necessary. It was not sufficient to restore availability.

I had to preserve evidence, restore verifiable application core files, and progressively separate the cost of executing WordPress from the cost of accepting and serving the incoming traffic.

The turning point was a controlled experiment: with PHP-FPM stopped and backend activity nearly absent, an Nginx worker still grew from a few megabytes to roughly 1.5 GB within seconds and was terminated by the kernel’s out-of-memory mechanism.

The final contingency separated the attacked route from the application runtime. Nginx served the homepage as precompressed static HTML, reducing its response payload from approximately 123 KB to approximately 16 KB for gzip-capable clients. Internal routes continued through WordPress.

Production stabilized while malicious traffic was still arriving.

That was the outcome: containment and operational recovery, not permanent DDoS protection or a claim that every historical component had been proven clean.

Environment and Constraints

The application ran on a small Linux virtual machine with a conventional layered web stack.

LayerEnvironment during the incident
Compute1 vCPU, approximately 2 GB RAM, approximately 1 GB swap
Front endNginx
Application executionApache and PHP-FPM
ApplicationWordPress and WooCommerce
Supporting servicesMySQL, hosting administration, and remote access on the same host

The original homepage request path was:

Internet → Nginx → Apache → PHP-FPM → WordPress / WooCommerce.

The constraint was not simply that the machine was small. Application execution, connection handling, logging, the database, and the ability to administer the machine all depended on the same limited resources.

A mitigation that exhausted the host could also remove the tools needed to recover it.

The emergency objective was therefore narrower than a platform modernization: recover a verifiable core, contain identified malicious artifacts, reduce the work triggered by the attacked route, and preserve enough operational headroom to keep the platform manageable.

I deliberately avoided mass upgrades during the incident. Changing WordPress and many plugins simultaneously would have introduced compatibility risks and made the effect of each intervention harder to interpret.

The Initial Symptom: Selective Casino Content

The malicious behavior did not reproduce consistently for every visitor.

That mattered. A normal-looking response from one browser could not establish that the application was behaving correctly for other requests.

The investigation found modified web-server rewrite rules that used the User-Agent to distinguish selected requests, including requests presenting a search-crawler-like identity. Those requests could be routed internally to a non-core PHP-named artifact associated with the casino content.

This was evidence of selective routing, or cloaking. It was not evidence that the requests actually came from a legitimate search crawler; a User-Agent is a request characteristic, not proof of identity.

The investigation also found remote code-loading mechanisms and an arbitrary-file upload capability. These findings established that the security problem extended beyond an unwanted page or a cosmetic content change.

Preserving Evidence Before Remediation

Restoring a working page and preserving the ability to understand the compromise were both requirements.

Before significant changes, I took a virtual-machine snapshot. Compromised files, pre-change backups, and a compressed copy of the affected webroot were preserved separately from the production application.

Identified artifacts were moved into quarantine outside the public webroot.

The sequence preserved a reviewable relationship between the state found, the intervention, and the resulting state. Removing suspicious files first and trying to reconstruct the evidence afterward would have broken that relationship.

The same principle later applied to the flood: preserve enough evidence of the pattern before reducing the logging that was threatening the server’s storage.

Confirmed Behavior Was Not the Same as Suspicion

The artifacts were not all assigned the same level of certainty.

The report identified mechanisms that retrieved external PHP content and passed it into execution, as well as a component capable of accepting arbitrary file uploads. The selected specimens corroborate those capabilities. Their presence does not establish every action that an attacker actually performed through them.

Other non-core or heavily obfuscated components were isolated because they were suspicious and increased risk. That did not justify inventing a complete behavioral analysis for each file.

Some plugins were also temporarily removed to reduce the exposed surface. Temporary isolation was a containment decision, not a finding that those plugins had caused the compromise.

Neither file ownership nor the presence of a vulnerable-looking component established the initial entry point. That question remained open.

Restoring a Verifiable WordPress Core

After preserving the compromised state, I obtained a clean official WordPress distribution matching the installed version and locale.

Core components were backed up and replaced from that trusted distribution. Non-core artifacts identified during the investigation were removed from the production webroot, and standard WordPress rewrite rules were restored.

The subsequent WP-CLI core checksum verification succeeded. The relevant boundary is precise: WP-CLI compares core files with the checksums published by WordPress.org.

This supported a claim of verifiable core integrity against the corresponding official distribution.

It did not certify every plugin, theme, uploaded file, account, or database record.

Additional containment included restricting PHP execution in uploads and removing unnecessary historical application content from the webroot. The available persistence and cron checks did not identify an obviously malicious cron task that explained persistence. That negative finding was limited to the checks performed during the response window.

Core recovery restored a known reference point. It did not close the security investigation.

The Second Incident: a Distributed HTTP Flood

While I was investigating and recovering the application, availability deteriorated across the stack.

The symptoms included HTTP 500 and 504 errors, TLS connection resets, PHP worker saturation, thousands of internal backend sockets, Nginx file-descriptor exhaustion, rapid log growth, and repeated OOM events.

These were related operational pressures, but they were not interchangeable diagnoses.

Measure the Request Pattern Before Choosing the Defense

Traffic samples showed approximately 2,000–2,500 HTTP requests per second during the higher-intensity measured intervals. The analysis of requested URLs showed that almost all of the flood was directed at /.

That range describes incident samples, not a sustained average across the entire event and not a benchmark of the server’s capacity.

The requests were distributed across many source addresses. Individual sources contributed comparatively modest amounts of traffic while the aggregate remained overwhelming.

Repetitive and anomalous User-Agent patterns provided additional signals for targeted filtering. They did not identify an actor, a country, or a relationship to the earlier compromise.

The practical finding was architectural: one application route was receiving most of the hostile load, but its execution path involved several services on the same machine.

Exhaustion Was Visible at More Than One Layer

Nginx reached an early limit of approximately 1,024 file descriptors. Apache accumulated thousands of internal sockets, and PHP-FPM reached its available workers.

I reduced the maximum number of PHP workers temporarily to prevent the application runtime from consuming all available memory.

This was a protection measure. It traded application concurrency for host survival; it did not create more processing capacity.

The next question was whether protecting PHP was enough.

Why the First Mitigations Were Insufficient

Blocking Individual Addresses Did Not Control the Aggregate

Specific blocks were useful against obvious sources or request patterns. They were not an adequate primary response to traffic distributed across many addresses.

The observed load was a property of the aggregate, not just the most active individual source.

Per-IP Rate Limiting Left the Same Aggregate Problem

A first global rate-limiting strategy reduced load, but it was not an acceptable policy for legitimate traffic. A later per-address limit did not resolve the incident.

With many participating addresses, each could remain within its own quota while their combined requests still exhausted the origin.

The lesson was not that rate limiting never works. It was that this per-IP policy, at this origin, did not adequately constrain the distributed workload being observed.

PHP-FPM Tuning Protected Only One Part of the Host

Reducing PHP concurrency addressed one source of memory pressure. Connections, front-end processing, responses, and logs could still consume resources even when the application performed little or no work.

Treating every availability failure as a WordPress performance problem would have left those costs unexplained.

Logs Became an Availability Problem

The flood also created a storage incident.

Logs grew from multiple gigabytes to tens of gigabytes, and the main partition reached full utilization.

After preserving sufficient evidence of the traffic pattern, I truncated the large flood-generated logs in a controlled manner to recover space. The relevant access logging was then disabled temporarily so that continuing traffic would not immediately consume the reclaimed storage.

A subsequent short measurement showed effectively no log growth over that test interval.

This was a deliberate observability trade-off, not a permanent logging strategy. The emergency response reduced one source of failure while sacrificing some request-level visibility. Bounded logging, retention, monitoring, and alerts remained follow-up work.

The Critical Discovery: PHP Wasn’t the Final Bottleneck

The kernel records showed repeated OOM terminations involving Nginx workers, with observed worker memory usage on the order of 1.0–1.65 GB.

On a host with approximately 2 GB of RAM, that left very little room for the database, operating system, application processes, and administration services.

I needed to distinguish application execution from front-end resource pressure.

The Controlled Experiment

I stopped PHP-FPM temporarily and reduced backend activity to almost zero. The incoming flood continued.

Under those conditions, an Nginx worker still grew rapidly and was terminated by OOM.

PHP-FPM stopped Backend activity nearly absent Flood still arriving
  1. Initially A few MB
  2. After ~2 seconds ~352 MB
  3. After ~4 seconds ~1.14 GB
  4. After ~6 seconds ~1.50 GB

Later observation: ~1.56 GB, followed by OOM termination.

Approximate snapshots recorded in the incident report, not a continuous memory trace or a synthetic load test. PHP execution had been removed from the experiment, but Nginx still exhausted memory.

The result changed the mitigation strategy.

There were now two separately observable problems: the expense of executing WordPress/WooCommerce under the flood, and the memory pressure that the arriving traffic could impose on Nginx without that execution.

This experiment did not identify a specific Nginx defect, memory leak, or protocol-level exploit. It established that application execution was not required for the failure observed at that point in the incident.

The response had to protect the host and reduce the cost of the attacked request path, rather than continue tuning PHP alone.

Emergency Resource Protection

I reduced Nginx concurrency substantially and applied aggressive timeouts to connections that were not progressing quickly.

The temporary worker_connections setting was 128. This is a per-worker connection setting, not a requests-per-second limit; Nginx also counts proxied-server connections within that limit.

The response also introduced systemd memory controls for the Nginx service. The incident report records MemoryHigh at 500M, MemoryMax at 600M, and MemorySwapMax at 100M.

These were emergency settings for this host, not a reusable sizing recommendation.

The intent was to contain the impact of another memory surge. Degrading or restarting the web service was preferable to allowing it to consume the host’s memory and take database access, remote administration, and other essential services with it.

There was a cost: lower concurrency and stricter limits could also reduce the service available to legitimate clients. The objective was controlled degradation while further mitigation was developed, not maximum throughput.

A Reload Was Not Always a Clean Resource Boundary

During the investigation, an old Nginx worker could remain alive draining connections after a reload while a new worker had already started.

One observation showed an old worker near 500 MB alongside a new worker using only about 9 MB. Looking only at the new configuration would have missed the memory still retained by the old process.

A clean restart removed the old worker and restored normal consumption in that observation. I then avoided unnecessary reloads while the flood remained active.

This was an incident-specific operational decision, not a general rule to replace graceful reloads with restarts.

A Static Homepage Was Only the First Step

Because almost all of the flood targeted /, that route offered a useful containment boundary.

I first made Nginx answer the homepage request with minimal static HTML. Those requests no longer triggered WordPress, WooCommerce, PHP, or Apache application execution.

PHP work, backend connections, and memory pressure dropped substantially. Internal pages remained on the dynamic application path.

The next objective was to recover the homepage’s presentation without putting the application back in the attacked request path.

Reconstruct, Validate, Then Expose to Load

An initial static capture had problems with rewritten URLs and asset names. I prepared a cleaner static version using the site’s existing production assets and tested it through a restricted preview route before exposing it at /.

Temporary plugin removal also had presentation consequences, including a shortcode that could appear as literal text. Restoring the visible homepage therefore required more than copying the first response into a file.

This production reconstruction was part of the incident response. None of that HTML, imagery, branding, or other client material is reproduced in this case study.

The Larger Static Page Still Cost Too Much

The reconstructed HTML was approximately 123 KB.

When it was served directly at the attacked route, Nginx memory consumption rose again. A worker approached 500 MB, and internal pages degraded.

I reverted the test to the minimal response.

That rollback mattered. It showed that removing PHP had reduced one cost, but had not made the homepage free to serve under sustained hostile traffic.

The larger-response test did not isolate a precise internal allocation mechanism. It showed that this response, in the active incident conditions, was still operationally too expensive.

Precompression Changed the Cost of the Response

The next iteration reduced the bytes sent for each gzip-capable homepage request without compressing the same HTML repeatedly at request time.

I generated a gzip-compressed version in advance and configured Nginx to serve it through gzip_static. Nginx’s static gzip module serves an existing precompressed file; the transformation does not need to be repeated for every response.

Reconstructed static HTML ~123 KB
Precompressed HTML response ~16 KB
Approximately 87% less HTML payload for the verified gzip-capable response. This is not a measurement of total page weight, total network traffic, or memory reduction.

External HTTP verification confirmed a successful HTTP/2 response, HTML content, gzip encoding, and a compressed content length of approximately 16 KB.

That verification applied to a request advertising gzip support. It did not establish that every client, or every hostile request, accepted or received the smaller representation.

Precompression was one part of the final contingency, alongside route isolation and resource limits. The report does not provide an isolated before-and-after benchmark that attributes all of the stabilization to compression alone.

Final Contingency Architecture

The resulting architecture kept the application available without requiring it to execute the route receiving most of the flood.

Implemented during the emergency

Internet Legitimate and malicious traffic
Nginx Reduced concurrency · timeouts · service memory limits

Route /

Precompressed static HTML ~16 KB when gzip is accepted

No WordPress or WooCommerce execution. No PHP or Apache application execution for this route.

Internal routes

  1. Apache
  2. PHP-FPM
  3. WordPress / WooCommerce

Internal pages continued through the dynamic application.

The homepage path was isolated, not the entire platform converted to static hosting. The origin still received hostile traffic; this was a temporary containment architecture, not an edge DDoS defense.

The important separation was between the attacked route and the application runtime.

This kept basic commercial functionality available while reducing the work demanded of the origin by the dominant request pattern.

The static homepage was also a temporary publishing constraint. By design, changes to the dynamic homepage would not automatically update its static representation. That is an architectural trade-off of this contingency, not an additional incident outcome measured in the report.

What Worked, What Didn’t, and Why

InterventionObserved result or limitation
Individual source and request-pattern blocksUseful for obvious matches; insufficient against the distributed aggregate.
Global origin rate limitReduced load; did not provide an appropriate policy for legitimate traffic.
Per-IP origin rate limitIndividual quotas did not adequately constrain combined load.
PHP-FPM concurrency reductionBounded PHP concurrency to protect host memory; did not resolve Nginx exhaustion.
Log containmentRecovered disk space and curtailed growth, at a cost to visibility.
Nginx connection and service memory limitsAdded emergency bounds intended to protect the host, with reduced service capacity.
Minimal static homepageReduced application execution and backend pressure.
Larger uncompressed static homepageMemory pressure returned and internal routes degraded; the test was reverted.
Precompressed static homepage with resource controlsFormed the final contingency under which operational stability was recovered.
Origin-only defense as a permanent architectureLeft the host exposed to direct traffic and future changes in the attack.

The unsuccessful or incomplete steps were not incidental to the story. They supplied the evidence needed to choose the next intervention.

What Stabilized and What Remained Open

At the end of the intervention, the WordPress core verified against official checksums. The malicious rewrite configuration had been preserved and replaced, identified malicious or suspicious artifacts had been quarantined, and evidence of the compromised state had been retained.

Log growth had been contained. PHP and Nginx had emergency resource protections. The homepage was served directly as precompressed static HTML, internal pages remained on WordPress, and the host had regained operational stability.

The report does not establish a sustained availability percentage, a latency improvement, an error-rate reduction, or an attack end time. None is inferred here.

Several risks remained:

  • The initial entry point and the full extent of historical compromise were unresolved.
  • Plugins, themes, uploads, and database content still required a deeper audit.
  • Administrative accounts, credentials, and infrastructure access required review and credential rotation.
  • The origin remained directly exposed to incoming traffic.
  • The contingency specifically addressed the route dominating the measured flood; it was not proof of protection for other routes or future attack patterns.
  • Logging, alerting, backups, recovery procedures, and capacity required follow-up work.

Containment created room for those tasks. It did not replace them.

The next phase should move distributed Layer 7 mitigation toward the edge, before unwanted requests consume the origin’s constrained resources.

Recommended next phase · not implemented during the incident

  1. Internet Incoming requests
  2. Edge DDoS protection / WAF Filtering, challenges, rate policies, and appropriate caching
  3. Origin firewall Restrict web ingress to authorized edge sources
  4. Nginx Bounded origin resources
  5. Application Apache / PHP-FPM / WordPress / WooCommerce
Separate traffic admission from application execution. Administrative access must remain explicitly controlled, and direct-origin bypass must be addressed.

A network firewall remains useful, but allowing web traffic on standard ports does not by itself distinguish a legitimate shopper from an automated client making valid HTTP requests.

The edge layer should apply application-aware policy, with protection for the homepage and other sensitive routes. The origin should then accept web ingress from authorized edge infrastructure, alongside the specifically required administrative access.

Changing DNS alone is not the entire solution. A reachable origin can bypass the intended protection; Cloudflare’s origin-security guidance describes restricting direct access.

Those controls need business validation. Payment integrations, legitimate crawlers, external services, monitoring, and real customers must continue to work. Broad geographic blocking should not be introduced without considering those dependencies.

This edge phase was intentionally deferred during the emergency. DNS changes, configuration, and commercial validation introduced a different set of risks and deserved a controlled implementation window.

It was a recommendation, not an architecture already deployed at incident closure.

Lessons Learned

Integrity and Availability Needed Separate Verification

A successful core checksum result answered an integrity question. It did not explain a full disk, file-descriptor exhaustion, or a front-end worker consuming most of the host’s memory.

Each failure mode needed its own observations and its own verification boundary.

Remove a Suspected Dependency and Observe What Remains

Stopping PHP temporarily made the investigation more precise. The remaining failure ruled out PHP execution as a necessary condition for that specific OOM event.

The experiment narrowed the explanation without pretending to establish a deeper root cause than the evidence supported.

A Smaller Execution Path Can Still Have a Significant Cost

Static delivery removed application work. The larger-response test showed that this alone was not enough under the observed flood.

The request path and the response representation both mattered.

Protect Recoverability Before Optimizing Throughput

Worker limits, memory controls, and log containment accepted reduced service capacity to preserve the machine and the ability to operate it.

That trade-off must be explicit. Emergency limits should not become permanent production policy without review.

Document the Boundary of the Result

The platform was stabilized. The attack had not been proven to have ended, and the entire application estate had not been forensically certified.

Clear outcome language is part of engineering rigor, especially when the intervention succeeds before the investigation is complete.

Incident Response Decision Records

These records summarize decisions documented in the report. Their identifiers and structure are retrospective editorial organization, not a claim that a formal decision log existed during the emergency.

Incident Response Decision Record

IR-001 — Preserve Evidence Before Remediation

Status
Applied
Priority
Preserve auditability

Context

Recovery would replace or remove parts of the compromised state.

Decision

Take a machine snapshot and preserve affected files and pre-change backups before significant remediation. Quarantine identified artifacts outside the public webroot.

Trade-off and Boundary

Evidence handling adds work during an emergency, but preserves the ability to review what was found and what changed. Preservation is not, by itself, a complete forensic investigation.

Incident Response Decision Record

IR-002 — Restore Application Core From Trusted Sources

Status
Applied and verified
Scope
WordPress core

Context

The application core needed a trustworthy reference without adding the compatibility uncertainty of a mass upgrade.

Decision

Restore core components from the official distribution matching the installed version and locale, then verify against official checksums.

Trade-off and Boundary

The intervention restored verifiable core files while deferring controlled upgrades. It did not validate every plugin, theme, upload, or database record.

Incident Response Decision Record

IR-003 — Separate the Attacked Route From the Application Runtime

Status
Applied temporarily
Scope
Homepage route

Context

Almost all measured flood requests targeted the homepage, whose original execution path invoked the dynamic application.

Decision

Serve the homepage directly through Nginx, first as minimal HTML and then as a reconstructed precompressed representation. Leave internal routes on the application path.

Trade-off and Boundary

The homepage became a separate static representation, while the origin still handled incoming traffic. The larger uncompressed trial was reverted when resource pressure returned.

Incident Response Decision Record

IR-004 — Bound Nginx Resource Consumption to Protect the Host

Status
Applied for contingency
Priority
Host recoverability

Context

Nginx workers could consume most of the machine's memory, even with PHP execution removed.

Decision

Reduce concurrency, tighten timeouts, and introduce service memory controls intended to contain front-end resource exhaustion.

Trade-off and Boundary

Accept reduced web-service capacity or service-level failure in preference to uncontrolled host-wide exhaustion. These incident-specific limits were not a permanent capacity plan.

Incident Response Decision Record

IR-005 — Prefer Measurements and Controlled Experiments Over Assumptions

Status
Applied throughout response
Turning point
PHP isolation experiment

Context

Application saturation was visible, but it did not explain all of the observed memory behavior.

Decision

Stop PHP-FPM temporarily, minimize backend activity, and observe Nginx under the continuing flood. Test larger static delivery and revert when it degraded the platform.

Trade-off and Boundary

Controlled interruption and rollback were part of the diagnosis. The observations supported changing mitigation priorities, not declaring a specific Nginx vulnerability or an attacker relationship.

Incident Response Decision Record

IR-006 — Move Distributed L7 Mitigation to the Edge

Status
Recommended; not implemented
Phase
Post-incident architecture

Context

Origin-only controls left the small host responsible for receiving the hostile workload.

Decision

Recommend an edge DDoS/WAF layer with origin firewall restrictions and explicit protection against direct-origin bypass.

Trade-off and Boundary

DNS, traffic policy, and integration validation required a separate implementation phase. The contingency was not presented as a substitute for that work.

Decision Summary

RecordEngineering boundaryState at incident closure
IR-001Preserve the state before changing itApplied
IR-002Verify core integrity without overstating scopeApplied and verified
IR-003Remove application work from the attacked routeTemporary contingency
IR-004Protect the host from web-service exhaustionEmergency controls
IR-005Let experiments change the diagnosisApplied
IR-006Move traffic admission ahead of the originRecommended next phase

Final Thoughts

The decisive change was not finding a single setting that made the attack disappear.

It was separating the problems carefully enough to act on each one: compromised application state, expensive dynamic execution, front-end memory pressure, storage exhaustion, and an origin that was receiving more hostile traffic than it could safely absorb.

Evidence preservation made recovery reviewable. Controlled experiments changed the diagnosis. Temporary architectural separation made stabilization possible.

The resulting system was stable enough to operate and continue the work. That was a meaningful engineering result, with clearly stated limits.

Explore all Engineering Case Studies or read ECS-001: Designing a Federated Enterprise Search Platform.


© 2024. All rights reserved.