Archtin
All articles
System DesignAIDevOpsArchitecture14 min read

Stop Adding AI to Everything: Choose the Simplest Reliable Solution

Six engineering scenarios where algorithms, rules, and automation beat an LLM—and how to recognize the problems where AI genuinely helps.

Your API starts returning errors. Someone proposes an agent to watch the logs. A query is slow, so another team suggests an AI database optimizer. A deployment fails, and the next proposal is an agent with production access.

These ideas sound modern. But they often introduce a probabilistic component before anyone has identified why the existing system is failing. The question is not whether AI is impressive. It is whether it is the right mechanism for this particular job.

The principle
Do not replace a well-defined algorithm, rule, or workflow with an LLM just because an LLM can produce an answer. Use the simplest solution that reliably satisfies the requirements. Sometimes that solution includes AI; often it does not.

AI is a tool—not a substitute for architecture

A token bucket answers a precise question: does this caller have enough request budget right now? A language model answers a different kind of question: given this context, what response is plausible? Those capabilities are not interchangeable.

For a fixed model version and decoding setup, you can reduce output variability. That does not make its judgment a formally specified policy. It can still misinterpret inputs, omit constraints, and change behavior when the prompt or evidence changes. Structured output constrains the answer’s shape, not its correctness.

Meanwhile, deterministic systems are not automatically correct. A bad alert threshold creates noise; an incorrectly implemented retry creates duplicate payments. Their advantage is that the behavior can be specified, tested, audited, and bounded without asking a model to reinterpret the policy on every request.

Production control pathRequestTyped inputPolicy + algorithmEnforce invariantsActionBounded outcomeAI investigation path — outside request enforcementLogs + metrics + traces → evidence → hypotheses → engineer reviewApproved actions still pass permission, policy, and validation checks.
Separate enforcement from interpretation. AI can assist the engineer without becoming the authority for every production decision.
RequirementUsually start withAI may help with
Exact decision under a tight deadlineAlgorithm or policy engineOffline analysis of policy effectiveness
Known response to a known failureTested automation or runbookExplaining why the automation ran
Ambiguous evidence across many sourcesObservability and investigationSummaries, hypotheses, and evidence retrieval
Interpretation of unstructured languageSearch, parsers, or rules where sufficientSemantic understanding and generation

1. Log monitoring: detect the condition with rules

Your service is returning too many 500 errors. The requirement is clear: notify the on-call engineer if the proportion of failed requests stays above a threshold. An LLM reading every log line is an expensive and less predictable way to calculate a ratio.

Build the observability path first

Instrument request counters and latency histograms. Prometheus evaluates metric-based alert rules; Alertmanager handles grouping, routing, silences, and notifications; Grafana displays the data. For searching actual log messages, use a log system such as Loki or Elasticsearch. Prometheus is not a general-purpose log-search engine.

1Service metrics2Prometheus rule3Alertmanager4On-call channel
Measure and alert deterministically. Investigating why the alert fired is a separate task.
sum(rate(http_requests_total{service="checkout",status=~"5.."}[5m]))
/
sum(rate(http_requests_total{service="checkout"}[5m]))
> 0.10

# In the alert rule, add a persistence condition such as:
# for: 5m
Illustrative PromQL: the selector and labels must match your instrumentation.

A ten-percent error ratio at two requests per minute means something different from the same ratio at ten thousand requests per second. Add a traffic floor where appropriate, distinguish routes and dependencies, and use SLO burn-rate alerts when that better reflects user impact. CPU above ninety percent is a capacity signal, not proof that users are affected.

  • Known condition: rising 5xx ratio, exhausted disk space, or a growing queue backlog. Use a rule.
  • Unknown cause: errors began after a deployment, but only in one region and only for one customer cohort. AI may help assemble and summarize the evidence.
  • Failure mode: sending all logs to a model adds sensitive-data exposure, processing cost, and a dependency that may fail during the incident itself.

The alert should still fire if the AI provider is unavailable. Investigation assistance is optional; detection of a known failure is not.

2. API rate limiting: use the right algorithm

“Can AI decide which requests to block?” bundles two problems together. Enforcing a published quota is an algorithmic problem. Identifying novel abusive behavior may be a statistical classification problem. Solve them separately.

For a policy of ten requests per second with short bursts up to twenty, a token bucket is a natural fit. It replenishes a bounded budget over time; each accepted request spends a token. A leaky-bucket-style scheduler is useful when the goal is to smooth outgoing work instead.

Refill: 10 tokens / secondCapacity: 20 tokens · stored state: 8 remainingToken available → acceptAtomically consume one tokenEmpty → 429 or bounded waitA policy decision, not a prediction
The quota is explicit. Every caller sees the same defined enforcement logic, independent of model interpretation.

Distributed enforcement needs atomic state

If many application instances share a quota, store the bucket state centrally or use a deliberately partitioned design. With Redis, replenish-and-consume must be one atomic operation—for example, a carefully implemented Lua script. A separate read followed by a write can allow concurrent requests to overspend the same tokens.

  • Choose the key: account, API key, tenant, endpoint, or a combination. IP-only policies can punish many users behind one network.
  • Define burst size, refill rate, clock assumptions, expiry, and multi-region behavior.
  • Return an explicit rejection such as HTTP 429 and suitable retry guidance when the quota is exhausted.
  • Decide what happens when the limiter is unavailable: fail open, fail closed, or use a bounded local fallback. Choose by the endpoint’s risk.

Risk scoring can inform separate abuse controls. It should not make the contractual quota mysterious, and an LLM does not need to sit in the latency-sensitive request path. For implementation trade-offs, see designing a rate limiter.

3. Database optimization: diagnose before automating

A slow API does not establish that the database needs AI—or even that the database is the bottleneck. The request may spend most of its time waiting for a connection, calling another service, or executing hundreds of individually fast queries.

1Trace request2Find slow SQL3Inspect plan4Change + measure
Find where time is spent, inspect the actual plan, change one thing, and measure the result.
  1. Locate the delay. Separate application work, network time, pool wait, query execution, and lock wait.
  2. Inspect the workload. Slow query logs and query statistics reveal repeated expensive statements and N+1 access patterns.
  3. Read the plan. Look at scans, joins, sorts, estimated versus actual row counts, and buffer activity.
  4. Make a targeted change. Rewrite the query, remove redundant round trips, update statistics, or add an index justified by the access pattern.
  5. Re-measure. Test representative parameters, dataset size, concurrency, write load, and tail latency.
EXPLAIN
SELECT id, created_at
FROM orders
WHERE account_id = 42
ORDER BY created_at DESC
LIMIT 50;

-- Evaluate this candidate against the real workload:
CREATE INDEX ON orders (account_id, created_at DESC);
An illustrative PostgreSQL query and candidate index—not a universal optimization.
EXPLAIN ANALYZE executes the statement
Plain EXPLAIN shows the planned execution. EXPLAIN ANALYZE actually runs it, including writes. Use a safe environment and representative data; even a read can be expensive on production. A transaction rollback does not undo every possible external side effect.

Indexes improve some reads but consume storage and add write maintenance. Connection pooling bounds concurrent connections; it does not make a bad query fast. Increasing the pool may simply move the queue into an already saturated database.

AI can explain a plan, draft candidate queries, or suggest hypotheses. Have an engineer validate those suggestions against real measurements before executing schema changes. Continue with a database at 100% CPU and Little’s Law for concurrency and pool sizing.

4. Cache invalidation: define freshness, not predictions

“Can AI predict what we should cache?” might be worth exploring at enormous scale with unusual access patterns. But it does not solve the more fundamental question: when is a cached value no longer safe to use?

Start with a freshness contract. A product description may tolerate sixty seconds of staleness. An authorization decision or inventory reservation may not. Those are business and correctness requirements, not things a model should infer from popularity.

1Read cache2Miss: query DB3Fill with TTL4Write: invalidate
Cache-aside defines the read path. The write path must separately define how old cached values stop being served.
MechanismWhat it controlsWhat it does not guarantee
TTLHow long an entry is retained before expiryImmediate freshness after a write
LRU / LFU evictionWhich keys to remove under memory pressureThat retained values are current
Explicit invalidation / versioned keysWhether old data remains eligible to serveNo race conditions without a designed protocol
Request coalescing / TTL jitterConcurrent refill pressure and synchronized expiryCorrectness of the underlying data

The race matters more than the prediction

Imagine a reader misses the cache and fetches an old database value. A writer commits a newer value and deletes the cache key. The reader then stores the old value after that deletion. “Invalidate on write” was implemented, yet stale data was reintroduced. Depending on the requirements, versions, conditional writes, event ordering, shorter expiry, or bypassing the cache may be necessary.

Also plan for cache stampedes, hot keys, and cache outages. Popularity prediction can improve hit rate, but it cannot establish correctness. Get the cache-aside protocol right before adding intelligence to it.

5. Service communication: contracts before agents

An order workflow charges a payment method and reserves inventory. If those steps are fixed, an agent deciding which service to call adds uncertainty without adding capability. Use a typed workflow or event-driven orchestration with explicit states.

1Create order2Reserve stock3Authorize payment4Confirm or undo
A known transaction follows explicit states. Failures and compensation are part of the design, not improvised by a model.

API contracts define valid inputs and outputs. Service discovery locates healthy instances; load balancing distributes requests; timeouts bound waits. None of those mechanisms needs a language model to reinterpret it.

  • Retries: retry only suitable transient failures, within a bounded budget, with backoff and jitter.
  • Idempotency: use a stable operation identifier so a retry does not charge a customer twice.
  • Circuit breakers: stop repeated calls to a failing dependency, rather than amplifying an outage.
  • Compensation: define what happens if stock is reserved but payment fails, or a response is lost after success.

An AI assistant may interpret a user’s free-form request and propose a structured action. The action still passes schema validation, authorization, amount limits, and an explicit state machine. Natural-language flexibility belongs at the input boundary, not in the invariants that protect money and inventory.

For these foundations, read the saga pattern, idempotency, and circuit breakers.

6. Deployment failures: automate known responses

A new version causes a measurable increase in errors. If the established response is to stop the rollout and restore the previous version, an AI agent is not required. A deployment controller can evaluate predefined health gates and execute a tested response.

1Small canary2Measure health3Pass: expand4Fail: halt / revert
This is a feedback loop with explicit gates. Investigation can happen in parallel without delaying a known safe response.

Make the response safe before making it automatic

Readiness checks decide whether an instance should receive traffic. Liveness checks decide whether it should be restarted. A canary compares a small new population against a baseline using user-relevant signals. A circuit breaker contains dependency failure; it is not itself a deployment rollback mechanism.

Rollback also has constraints. An old application binary may no longer understand a changed database schema. A reverted service cannot necessarily reverse messages already emitted or payments already processed. Use backward-compatible schema evolution and define recovery paths for irreversible side effects.

  • Specify the signal, observation window, minimum traffic, and abort threshold.
  • Version the runbook and test it under realistic failure conditions.
  • Limit automated actions, require cooldowns, and prevent remediation loops.
  • Record what happened and preserve a manual stop mechanism.

AI becomes useful when the predefined response is insufficient: perhaps latency rises without an error spike, only certain pods fail, or a dependency changes behavior. Let it assist investigation, not invent unrestricted production changes. See liveness, readiness, and delivery misconceptions and the local-to-production journey.

Where AI actually helps: reasoning over messy evidence

Now consider an incident with logs, metrics, traces, deployment history, Kubernetes events, and database errors. The engineer’s question is not “is the error ratio above ten percent?” It is “what probably changed, and which explanation best fits these observations?”

This is a more plausible AI use case because the inputs are heterogeneous and the investigation is partly linguistic. An assistant can retrieve relevant records, summarize a timeline, compare hypotheses, and suggest the next check. Its output is an aid to reasoning—not proof of causation.

1Collect evidence2Build timeline3Rank hypotheses4Engineer verifies
An incident copilot should produce evidence-linked hypotheses, not a confident story without supporting observations.

A useful incident answer separates facts from guesses

Observation:
  Checkout p95 increased after rollout v42.
  Database pool wait increased; SQL execution time stayed stable.

Hypothesis:
  The new release holds connections longer than before.

Evidence:
  Links to the rollout event, pool metric, and representative traces.

Next check:
  Compare connection lifetime and transaction boundaries in v41/v42.

Uncertainty:
  Timing is correlated; the deployment is not yet proven causal.
Illustrative response format for an incident assistant.

The next check is valuable because it is falsifiable. A claim such as “the database is overloaded” without evidence, scope, or an alternative explanation is not much help.

Start read-only
Give the assistant narrowly scoped retrieval access, redact secrets and personal data, and treat logs and documents as untrusted input. A malicious log line must not become an instruction to execute a command. Suggested actions need independent authorization and validation.

If you later permit actions, use allowlisted tools, restricted credentials, bounded impact, approvals for risky changes, and an audit trail. Test those controls independently of the prompt. An instruction saying “be careful” is not a security boundary.

Measure whether the assistant actually helps: time to find useful evidence, unsupported-claim frequency, relevance of proposed checks, operator acceptance, and total investigation time. Evaluate on held-out incidents and compare against search plus existing runbooks. For a practical build, see the AI incident copilot project guide.

A decision framework before you add a model

  1. Write down the decision. “Block a caller after its budget is exhausted” is precise. “Improve reliability” is not.
  2. Define correctness. Which answers are valid? What are the cost and consequences of a false positive or false negative?
  3. Build the non-AI baseline. Try an algorithm, rule, parser, database query, retrieval system, or deterministic workflow.
  4. Locate the uncertainty. Is the hard part unstructured language, incomplete evidence, or changing patterns—or simply missing instrumentation?
  5. Budget latency and failure. What happens when the model is slow, unavailable, wrong, or more expensive than expected?
  6. Evaluate the improvement. Use representative examples, explicit scoring, and a realistic baseline—not a handful of impressive demos.
  7. Contain the authority. Keep critical invariants enforced outside the model and provide a fallback or manual path.

Specialized machine-learning systems can legitimately support anomaly detection, fraud scoring, predictive caching, or workload forecasting. That does not mean every one of those tasks needs a general-purpose LLM or autonomous agent. “AI versus rules” is not a universal binary; the right architecture may combine deterministic enforcement with a narrowly evaluated learned signal.

ScenarioReliable first choiceA sensible AI role
Known error thresholdMetric alert + notification routingSummarize the incident and retrieve context
Published API quotaToken bucket / equivalent limiterAssist offline abuse analysis
Slow database queryTrace + plan + measured optimizationExplain plans and propose candidates
Predictable data accessFreshness policy + cache strategyExplore prefetching when a baseline is insufficient
Known service workflowContracts + state machine + resilienceTranslate language into validated commands
Known rollout regressionCanary gates + safe halt / rollbackInvestigate ambiguous failure evidence

Ask what the system needs—not where AI can fit

The practical lesson is not “never use AI.” It is “do not use AI to avoid understanding the engineering problem.” An agent cannot compensate for missing metrics, unclear cache semantics, unsafe retries, or an untested rollback plan.

  • Use algorithms for precisely defined computations and budgets.
  • Use rules for known conditions with explicit responses.
  • Use observability to make the system’s behavior visible.
  • Use tested workflows to enforce contracts and invariants.
  • Use AI when interpreting language or uncertain evidence adds measurable value.
  • Keep permissions, correctness checks, and blast-radius limits outside the model.
An interview-ready answer
“I would keep this control path deterministic because the policy is exact and failure is costly. I would use AI off the critical path to summarize evidence or suggest hypotheses, and only expand its role after comparing it with a simpler baseline.”

Good engineers do not begin with “where can we add AI?” They begin with “what is the simplest reliable solution?” Sometimes the answer is AI. Sometimes it is a counter, an index, a timeout, or a well-written runbook.

Further reading

Keep reading

Suggested next articles based on this one.

Design it, don't just read it.

Practise LLD and system design problems with structured rubrics and AI feedback.

Start practising free