Best Practices for Observability in Microservices
Standardise telemetry, alert on user impact, trace critical flows and control telemetry volume to get observability working at scale.
Read moreBlog posts in the Performance category
Standardise telemetry, alert on user impact, trace critical flows and control telemetry volume to get observability working at scale.
Read moreMeasure 2-4 weeks of usage, set requests and limits from p90–p99, and test under live traffic to cut costs and avoid throttling or OOMs.
Read morePublic cloud often uses less energy per workload; hybrid can be greener when latency, data residency or placement matter.
Read moreUse forecast-led predictive scaling with reactive fallback to reduce latency and cloud spend for repeatable workloads.
Read moreCut serverless tail latency by pre-warming; size provisioned concurrency for p95/p99, test lighter runtimes and schedule to save cost.
Read moreCut search costs and improve reliability: keep shard sizes 10–50 GB, match replicas to failure needs, tier old data and review sizing regularly.
Read moreScore workloads by cost, latency, data locality and compliance to place them on‑prem, private, public or edge, and review placement regularly.
Read moreUse user-focused SLIs, set SLO targets and manage error budgets to guide releases, scaling and cloud spend.
Read morePrioritise p99 latency, errors, throttles, concurrency, memory and cost in CloudWatch; use X‑Ray for tracing and tune memory, timeout and concurrency.
Read moreMeasure service mesh CPU, memory and telemetry overhead, convert deltas to monthly £ and reduce costs by scope, proxy and telemetry tuning.
Read moreUse minimal CDN geolocation: prioritise compliance and availability, limit cache variants, and measure routing accuracy.
Read morePipelines — not tools — are the real limit: standardise templates, cut queue time, automate policy checks and shift ownership to scale enterprise CI/CD.
Read more