Stop Adding AI to Everything: Choose the Simplest Reliable Solution
Six engineering scenarios where algorithms, rules, and automation beat an LLM—and how to recognize the problems where AI genuinely helps.
Your API starts returning errors. Someone proposes an agent to watch the logs. A query is slow, so another team suggests an AI database optimizer. A deployment fails, and the next proposal is an agent with production access.
These ideas sound modern. But they often introduce a probabilistic component before anyone has identified why the existing system is failing. The question is not whether AI is impressive. It is whether it is the right mechanism for this particular job.
AI is a tool—not a substitute for architecture
A token bucket answers a precise question: does this caller have enough request budget right now? A language model answers a different kind of question: given this context, what response is plausible? Those capabilities are not interchangeable.
For a fixed model version and decoding setup, you can reduce output variability. That does not make its judgment a formally specified policy. It can still misinterpret inputs, omit constraints, and change behavior when the prompt or evidence changes. Structured output constrains the answer’s shape, not its correctness.
Meanwhile, deterministic systems are not automatically correct. A bad alert threshold creates noise; an incorrectly implemented retry creates duplicate payments. Their advantage is that the behavior can be specified, tested, audited, and bounded without asking a model to reinterpret the policy on every request.
| Requirement | Usually start with | AI may help with |
|---|---|---|
| Exact decision under a tight deadline | Algorithm or policy engine | Offline analysis of policy effectiveness |
| Known response to a known failure | Tested automation or runbook | Explaining why the automation ran |
| Ambiguous evidence across many sources | Observability and investigation | Summaries, hypotheses, and evidence retrieval |
| Interpretation of unstructured language | Search, parsers, or rules where sufficient | Semantic understanding and generation |
1. Log monitoring: detect the condition with rules
Your service is returning too many 500 errors. The requirement is clear: notify the on-call engineer if the proportion of failed requests stays above a threshold. An LLM reading every log line is an expensive and less predictable way to calculate a ratio.
Build the observability path first
Instrument request counters and latency histograms. Prometheus evaluates metric-based alert rules; Alertmanager handles grouping, routing, silences, and notifications; Grafana displays the data. For searching actual log messages, use a log system such as Loki or Elasticsearch. Prometheus is not a general-purpose log-search engine.
sum(rate(http_requests_total{service="checkout",status=~"5.."}[5m]))
/
sum(rate(http_requests_total{service="checkout"}[5m]))
> 0.10
# In the alert rule, add a persistence condition such as:
# for: 5mA ten-percent error ratio at two requests per minute means something different from the same ratio at ten thousand requests per second. Add a traffic floor where appropriate, distinguish routes and dependencies, and use SLO burn-rate alerts when that better reflects user impact. CPU above ninety percent is a capacity signal, not proof that users are affected.
- Known condition: rising 5xx ratio, exhausted disk space, or a growing queue backlog. Use a rule.
- Unknown cause: errors began after a deployment, but only in one region and only for one customer cohort. AI may help assemble and summarize the evidence.
- Failure mode: sending all logs to a model adds sensitive-data exposure, processing cost, and a dependency that may fail during the incident itself.
The alert should still fire if the AI provider is unavailable. Investigation assistance is optional; detection of a known failure is not.
2. API rate limiting: use the right algorithm
“Can AI decide which requests to block?” bundles two problems together. Enforcing a published quota is an algorithmic problem. Identifying novel abusive behavior may be a statistical classification problem. Solve them separately.
For a policy of ten requests per second with short bursts up to twenty, a token bucket is a natural fit. It replenishes a bounded budget over time; each accepted request spends a token. A leaky-bucket-style scheduler is useful when the goal is to smooth outgoing work instead.
Distributed enforcement needs atomic state
If many application instances share a quota, store the bucket state centrally or use a deliberately partitioned design. With Redis, replenish-and-consume must be one atomic operation—for example, a carefully implemented Lua script. A separate read followed by a write can allow concurrent requests to overspend the same tokens.
- Choose the key: account, API key, tenant, endpoint, or a combination. IP-only policies can punish many users behind one network.
- Define burst size, refill rate, clock assumptions, expiry, and multi-region behavior.
- Return an explicit rejection such as HTTP 429 and suitable retry guidance when the quota is exhausted.
- Decide what happens when the limiter is unavailable: fail open, fail closed, or use a bounded local fallback. Choose by the endpoint’s risk.
Risk scoring can inform separate abuse controls. It should not make the contractual quota mysterious, and an LLM does not need to sit in the latency-sensitive request path. For implementation trade-offs, see designing a rate limiter.
3. Database optimization: diagnose before automating
A slow API does not establish that the database needs AI—or even that the database is the bottleneck. The request may spend most of its time waiting for a connection, calling another service, or executing hundreds of individually fast queries.
- Locate the delay. Separate application work, network time, pool wait, query execution, and lock wait.
- Inspect the workload. Slow query logs and query statistics reveal repeated expensive statements and N+1 access patterns.
- Read the plan. Look at scans, joins, sorts, estimated versus actual row counts, and buffer activity.
- Make a targeted change. Rewrite the query, remove redundant round trips, update statistics, or add an index justified by the access pattern.
- Re-measure. Test representative parameters, dataset size, concurrency, write load, and tail latency.
EXPLAIN SELECT id, created_at FROM orders WHERE account_id = 42 ORDER BY created_at DESC LIMIT 50; -- Evaluate this candidate against the real workload: CREATE INDEX ON orders (account_id, created_at DESC);
Indexes improve some reads but consume storage and add write maintenance. Connection pooling bounds concurrent connections; it does not make a bad query fast. Increasing the pool may simply move the queue into an already saturated database.
AI can explain a plan, draft candidate queries, or suggest hypotheses. Have an engineer validate those suggestions against real measurements before executing schema changes. Continue with a database at 100% CPU and Little’s Law for concurrency and pool sizing.
4. Cache invalidation: define freshness, not predictions
“Can AI predict what we should cache?” might be worth exploring at enormous scale with unusual access patterns. But it does not solve the more fundamental question: when is a cached value no longer safe to use?
Start with a freshness contract. A product description may tolerate sixty seconds of staleness. An authorization decision or inventory reservation may not. Those are business and correctness requirements, not things a model should infer from popularity.
| Mechanism | What it controls | What it does not guarantee |
|---|---|---|
| TTL | How long an entry is retained before expiry | Immediate freshness after a write |
| LRU / LFU eviction | Which keys to remove under memory pressure | That retained values are current |
| Explicit invalidation / versioned keys | Whether old data remains eligible to serve | No race conditions without a designed protocol |
| Request coalescing / TTL jitter | Concurrent refill pressure and synchronized expiry | Correctness of the underlying data |
The race matters more than the prediction
Imagine a reader misses the cache and fetches an old database value. A writer commits a newer value and deletes the cache key. The reader then stores the old value after that deletion. “Invalidate on write” was implemented, yet stale data was reintroduced. Depending on the requirements, versions, conditional writes, event ordering, shorter expiry, or bypassing the cache may be necessary.
Also plan for cache stampedes, hot keys, and cache outages. Popularity prediction can improve hit rate, but it cannot establish correctness. Get the cache-aside protocol right before adding intelligence to it.
5. Service communication: contracts before agents
An order workflow charges a payment method and reserves inventory. If those steps are fixed, an agent deciding which service to call adds uncertainty without adding capability. Use a typed workflow or event-driven orchestration with explicit states.
API contracts define valid inputs and outputs. Service discovery locates healthy instances; load balancing distributes requests; timeouts bound waits. None of those mechanisms needs a language model to reinterpret it.
- Retries: retry only suitable transient failures, within a bounded budget, with backoff and jitter.
- Idempotency: use a stable operation identifier so a retry does not charge a customer twice.
- Circuit breakers: stop repeated calls to a failing dependency, rather than amplifying an outage.
- Compensation: define what happens if stock is reserved but payment fails, or a response is lost after success.
An AI assistant may interpret a user’s free-form request and propose a structured action. The action still passes schema validation, authorization, amount limits, and an explicit state machine. Natural-language flexibility belongs at the input boundary, not in the invariants that protect money and inventory.
For these foundations, read the saga pattern, idempotency, and circuit breakers.
6. Deployment failures: automate known responses
A new version causes a measurable increase in errors. If the established response is to stop the rollout and restore the previous version, an AI agent is not required. A deployment controller can evaluate predefined health gates and execute a tested response.
Make the response safe before making it automatic
Readiness checks decide whether an instance should receive traffic. Liveness checks decide whether it should be restarted. A canary compares a small new population against a baseline using user-relevant signals. A circuit breaker contains dependency failure; it is not itself a deployment rollback mechanism.
Rollback also has constraints. An old application binary may no longer understand a changed database schema. A reverted service cannot necessarily reverse messages already emitted or payments already processed. Use backward-compatible schema evolution and define recovery paths for irreversible side effects.
- Specify the signal, observation window, minimum traffic, and abort threshold.
- Version the runbook and test it under realistic failure conditions.
- Limit automated actions, require cooldowns, and prevent remediation loops.
- Record what happened and preserve a manual stop mechanism.
AI becomes useful when the predefined response is insufficient: perhaps latency rises without an error spike, only certain pods fail, or a dependency changes behavior. Let it assist investigation, not invent unrestricted production changes. See liveness, readiness, and delivery misconceptions and the local-to-production journey.
Where AI actually helps: reasoning over messy evidence
Now consider an incident with logs, metrics, traces, deployment history, Kubernetes events, and database errors. The engineer’s question is not “is the error ratio above ten percent?” It is “what probably changed, and which explanation best fits these observations?”
This is a more plausible AI use case because the inputs are heterogeneous and the investigation is partly linguistic. An assistant can retrieve relevant records, summarize a timeline, compare hypotheses, and suggest the next check. Its output is an aid to reasoning—not proof of causation.
A useful incident answer separates facts from guesses
Observation: Checkout p95 increased after rollout v42. Database pool wait increased; SQL execution time stayed stable. Hypothesis: The new release holds connections longer than before. Evidence: Links to the rollout event, pool metric, and representative traces. Next check: Compare connection lifetime and transaction boundaries in v41/v42. Uncertainty: Timing is correlated; the deployment is not yet proven causal.
The next check is valuable because it is falsifiable. A claim such as “the database is overloaded” without evidence, scope, or an alternative explanation is not much help.
If you later permit actions, use allowlisted tools, restricted credentials, bounded impact, approvals for risky changes, and an audit trail. Test those controls independently of the prompt. An instruction saying “be careful” is not a security boundary.
Measure whether the assistant actually helps: time to find useful evidence, unsupported-claim frequency, relevance of proposed checks, operator acceptance, and total investigation time. Evaluate on held-out incidents and compare against search plus existing runbooks. For a practical build, see the AI incident copilot project guide.
A decision framework before you add a model
- Write down the decision. “Block a caller after its budget is exhausted” is precise. “Improve reliability” is not.
- Define correctness. Which answers are valid? What are the cost and consequences of a false positive or false negative?
- Build the non-AI baseline. Try an algorithm, rule, parser, database query, retrieval system, or deterministic workflow.
- Locate the uncertainty. Is the hard part unstructured language, incomplete evidence, or changing patterns—or simply missing instrumentation?
- Budget latency and failure. What happens when the model is slow, unavailable, wrong, or more expensive than expected?
- Evaluate the improvement. Use representative examples, explicit scoring, and a realistic baseline—not a handful of impressive demos.
- Contain the authority. Keep critical invariants enforced outside the model and provide a fallback or manual path.
Specialized machine-learning systems can legitimately support anomaly detection, fraud scoring, predictive caching, or workload forecasting. That does not mean every one of those tasks needs a general-purpose LLM or autonomous agent. “AI versus rules” is not a universal binary; the right architecture may combine deterministic enforcement with a narrowly evaluated learned signal.
| Scenario | Reliable first choice | A sensible AI role |
|---|---|---|
| Known error threshold | Metric alert + notification routing | Summarize the incident and retrieve context |
| Published API quota | Token bucket / equivalent limiter | Assist offline abuse analysis |
| Slow database query | Trace + plan + measured optimization | Explain plans and propose candidates |
| Predictable data access | Freshness policy + cache strategy | Explore prefetching when a baseline is insufficient |
| Known service workflow | Contracts + state machine + resilience | Translate language into validated commands |
| Known rollout regression | Canary gates + safe halt / rollback | Investigate ambiguous failure evidence |
Ask what the system needs—not where AI can fit
The practical lesson is not “never use AI.” It is “do not use AI to avoid understanding the engineering problem.” An agent cannot compensate for missing metrics, unclear cache semantics, unsafe retries, or an untested rollback plan.
- Use algorithms for precisely defined computations and budgets.
- Use rules for known conditions with explicit responses.
- Use observability to make the system’s behavior visible.
- Use tested workflows to enforce contracts and invariants.
- Use AI when interpreting language or uncertain evidence adds measurable value.
- Keep permissions, correctness checks, and blast-radius limits outside the model.
Good engineers do not begin with “where can we add AI?” They begin with “what is the simplest reliable solution?” Sometimes the answer is AI. Sometimes it is a counter, an index, a timeout, or a well-written runbook.
Further reading
- Prometheus alerting overview
- PostgreSQL: using EXPLAIN
- Redis key eviction policies
- Kubernetes Deployments and rollback behavior
Keep reading
Suggested next articles based on this one.
5 DevOps + AI + Cloud Projects That Actually Stand Out
Not another CRUD app. Five projects—incident copilot, AI CI/CD, self-healing Kubernetes, cost optimizer and security engine—with build guides and reference repos.
Read5 DevOps Concepts Developers Usually Get Wrong
Five familiar DevOps terms hide important operational differences. Learn the mechanism, the common mistake, and what changes in production.
ReadHow Code Flows: From Local Development to Production
A visual, end-to-end trace of how one code change becomes a versioned production deployment—and how the platform validates, routes, observes, scales and rolls it back.
ReadDesign it, don't just read it.
Practise LLD and system design problems with structured rubrics and AI feedback.
Start practising free