Evidence-Led Performance Profiling Experiment Planner
Guide Codex to build controlled profiling experiments, rank bottleneck hypotheses, and define evidence-based optimization, verification, and rollback decisions.
Design a measurement-first performance investigation for the supplied workload. Use available code, architecture, telemetry, profiles, logs, query plans, and benchmark results to isolate bottlenecks before proposing changes. ## Investigation context Performance symptom: [Performance symptom] Affected workload: [Affected workload] Baseline evidence: [Baseline evidence] Workload model: [Workload model] Code and architecture evidence: [Code and architecture evidence] Runtime and infrastructure: [Runtime and infrastructure] Available profiling tools: [Available profiling tools] Test environment: [Test environment] Acceptance thresholds: [Acceptance thresholds] Risk and authority constraints: [Risk and authority constraints] ## Input and access checks Treat these as blocking prerequisites for a conclusive experiment plan: - A precise workload boundary, such as an endpoint, query, queue worker, scheduled command, rendering path, or batch job. - A defined performance symptom and at least one measurable outcome, such as latency, throughput, queue lag, CPU time, allocation rate, memory growth, database time, or error rate. - A representative workload model or an explicit statement that representativeness is unknown. - A safe environment in which the proposed measurements can run, or existing execution evidence suitable for offline analysis. - Acceptance thresholds and authority limits for tests that may create load, execute queries, alter caches, or affect shared resources. Useful but non-blocking context includes deployment history, traces, flamegraphs, slow-query samples, database statistics, data-volume distributions, cache telemetry, queue metrics, dependency service-level objectives, and prior failed optimization attempts. If a blocking input is absent or conflicting, ask focused clarification questions first. Continue only with bounded work that remains valid without the missing input, label the resulting plan provisional, and preserve unknown values rather than estimating them. Reconcile conflicts between dashboards, logs, traces, and user reports by recording source, time window, aggregation, sampling, and environment differences. ## Codex operating boundaries Codex may inspect only files, repository content, logs, telemetry exports, and runtime facilities actually supplied or accessible in the current session. It may propose commands and code changes. It may execute read-only inspection or benchmark commands only when the environment provides execution access and the stated authority permits them. Do not claim access to production telemetry, profilers, databases, cloud consoles, or deployment systems unless access is demonstrated. Do not run production load tests, destructive database operations, cache flushes, data mutations, deployments, configuration changes, or infrastructure scaling actions. Do not use live customer records in test fixtures or expose secrets, tokens, personal data, payment data, or sensitive query parameters in output. Require explicit human authorization before any test against production or a shared environment; any use of EXPLAIN ANALYZE or another command that executes a query; any profiler with material overhead; cache invalidation; schema or index changes; concurrency increases; deployment; or changes involving payments, authorization, customer data, reporting accuracy, or production infrastructure. Stop if error rate, saturation, cost, lock time, queue growth, data integrity, or user impact crosses the supplied safety limit. ## Evidence discipline Maintain an evidence ledger using these states: - Supplied fact: directly present in the inputs. - Observed result: produced by an authorized command or experiment in this session and accompanied by its command, environment, timestamp or run identifier, and output reference. - Hypothesis: a testable explanation awaiting evidence. - Assumption: a bounded premise required to design the plan. - Unknown: information that cannot currently be established. - Conflict: incompatible evidence requiring reconciliation. Never describe a suspected bottleneck as confirmed. Never describe an optimization as tested, improved, verified, deployed, or safe unless the corresponding action occurred and evidence is available. Keep planned, executed, blocked, inconclusive, and verified work distinct. ## Investigation workflow ### 1. Normalize the performance question Define the workload boundary, affected users or downstream systems, symptom, environment, relevant time window, current baseline, target, and operational risk. Separate latency distributions from averages. Identify whether the issue concerns response time, throughput, tail latency, queue delay, resource consumption, scalability, or degradation over time. ### 2. Assess benchmark validity Translate the workload model into arrival rate, concurrency, request or job mix, payload and result-size distributions, data volume and cardinality, cache state, authentication state, tenant distribution, think time, dependency behavior, and run duration where relevant. Identify threats to validity, including: - Non-representative fixtures or database statistics. - Debug mode, tracing, logging, profiler, or coverage overhead. - Cold-start, just-in-time compilation, connection establishment, autoscaling, or warm-up effects. - Cold-cache and warm-cache results being mixed. - Background traffic, scheduled jobs, noisy neighbors, throttling, retries, or rate limits. - Open versus closed load-model mismatch and coordinated omission in latency measurement. - Client or load-generator saturation being mistaken for server saturation. - Different builds, configuration, feature flags, dependency versions, hardware, or data snapshots. - Too few repetitions, unstable variance, outliers without explanation, or time windows that hide tail behavior. Specify warm-up, ramp-up, steady-state, cool-down, repetition count, randomization or run ordering, and evidence needed to compare runs. Prefer identical build, configuration, data, and infrastructure for before-and-after comparisons. If statistical confidence cannot be estimated, report run count, spread, and uncertainty instead of asserting significance. ### 3. Map the critical path and resource model Trace the workload through applicable layers: client or load generator, web server, application middleware, controller or handler, serialization, database, cache, queue, file or network I/O, external APIs, and operating-system or container resources. Relate latency and throughput to utilization, saturation, and errors. For queued work, compare arrival rate with service rate and inspect queue depth, oldest-job age, retries, timeout behavior, worker concurrency, and downstream capacity. For databases, inspect query count, cumulative query time, plan shape, row estimates versus actual rows when safely available, scans, joins, sorts, temporary storage, lock waits, connection-pool pressure, and index selectivity. For caches, inspect hit ratio by operation, key cardinality, expiry behavior, stampede risk, eviction, serialization cost, and freshness requirements. ### 4. Build and rank bottleneck hypotheses Create testable hypotheses grounded in the evidence ledger. Consider only applicable mechanisms, such as N+1 access, high-cardinality query patterns, stale database statistics, missing or poorly ordered indexes, parameter-sensitive plans, lock contention, connection-pool exhaustion, excessive allocation or garbage collection, synchronous external calls, repeated serialization, oversized payloads, cache churn, cache stampedes, queue backpressure, retry amplification, thread or event-loop blocking, filesystem latency, CPU throttling, memory pressure, or load-generator limits. Rank each hypothesis by evidence strength, expected impact, likelihood, cost to test, experiment risk, and ability to isolate the cause. Explain competing explanations and what observation would distinguish them. ### 5. Design controlled profiling experiments For each hypothesis, specify: - Experiment identifier and hypothesis. - Evidence supporting and contradicting it. - Independent variable and controlled conditions. - Environment, dataset, cache state, workload phase, concurrency, duration, and repetitions. - Profiler, trace, metric, log, query-plan inspection, or benchmark mechanism. - Exact command or procedure only when supported by the supplied stack; otherwise mark it optional and name the prerequisite. - Expected observation if confirmed and expected observation if rejected. - Primary metric, guardrail metrics, units, aggregation, and collection location. - Observer-effect risk and a lower-overhead alternative. - Safety limit, stop condition, cleanup requirement, and required authorization. - Interpretation rule, confounders, and the next action for confirmed, rejected, or inconclusive outcomes. Prefer experiments that change one meaningful factor at a time. Pair wall-clock timing with layer-specific evidence such as traces, profiles, query plans, database waits, cache events, queue telemetry, or resource counters. Do not infer causation from correlation alone. ### 6. Select investigation commands safely Recommend only commands and tools compatible with the supplied runtime and available tooling. For every command, state purpose, target environment, read-only or mutating status, expected artifact, overhead, authorization requirement, and redaction needs. Distinguish a non-executing query-plan inspection from plan analysis that executes the query. Warn about table scans, locks, expensive aggregation, profiler overhead, large trace volumes, log amplification, and load generation. Where direct execution is unavailable, provide a runnable procedure for an authorized operator rather than fabricating output. ### 7. Gate optimization candidates on evidence Propose a candidate only when it is tied to a hypothesis and confirmation criterion. Evaluate local speedups against system-level trade-offs. Examples include read-versus-write cost for indexes, freshness and invalidation complexity for caching, memory-versus-CPU trade-offs, batching-versus-tail latency, concurrency-versus-downstream saturation, payload reduction-versus-client compatibility, and asynchronous work-versus delivery guarantees. For each candidate, specify the confirmed bottleneck required, smallest safe change, expected mechanism, affected components, correctness risks, operational risks, migration or compatibility concerns, expected metric movement, guardrail metrics, test strategy, rollout boundary, observability requirement, and rollback trigger. Do not recommend broad rewrites, scaling, caching, denormalization, or new infrastructure merely because they might improve performance. ### 8. Define verification and acceptance Require a before-and-after comparison under equivalent conditions. Record expected and actual observations for latency percentiles, throughput, errors, resource use, query behavior, queue health, cache behavior, dependency impact, and any workload-specific metric. Verify correctness alongside speed: response schema and semantics, result ordering and pagination, query-result accuracy, authorization and tenant isolation, cache freshness, idempotency, transaction behavior, duplicate or lost job handling, retry behavior, timeout paths, reporting totals, and representative edge cases. Classify each acceptance criterion as passed, failed, blocked, or inconclusive, with an evidence reference. A performance gain is not acceptable if correctness fails, errors exceed the threshold, resource use is displaced to another constrained component, tail latency regresses outside tolerance, or the result depends on an unrepresentative workload. ### 9. Prepare implementation and operational handoff Sequence work as baseline capture, low-overhead observation, isolating experiment, evidence review, smallest justified change, automated correctness checks, controlled benchmark, staged rollout, monitoring window, and rollback decision. Identify the human approver for consequential actions. Include rollback feasibility, recovery steps, and post-rollback verification. If no hypothesis is confirmed, recommend the next discriminating experiment rather than an optimization. ## Required deliverable Produce these sections: ### Performance Question and Input Status State the workload boundary, symptom, affected parties, supplied facts, blocking gaps, assumptions, conflicts, and whether the plan is ready, provisional, or blocked. ### Workload and Benchmark Validity Specification Provide the workload dimensions, environment controls, run phases, cache conditions, repetitions, comparability rules, validity threats, and load-generator capacity checks. ### Critical-Path and Resource Map Show components in execution order with observed or unknown latency contribution, throughput, utilization, saturation, errors, dependencies, and available evidence. ### Baseline Measurement Matrix Use columns: metric; unit and aggregation; source; capture method; workload phase; current value; target; regression boundary; guardrail status; evidence state. Include only applicable metrics and preserve unknown values. ### Evidence Ledger Use columns: identifier; statement; state; source or command; environment and time window; confidence; conflict or limitation; required follow-up. ### Ranked Bottleneck Hypotheses Use columns: rank; mechanism; supporting evidence; contradicting evidence; competing explanation; expected impact; test cost; test risk; discriminating observation; status. ### Profiling Experiment Cards Create one card per experiment containing every field required in the controlled experiment step, including authorization, stop conditions, confounders, and confirmed, rejected, or inconclusive branches. ### Command and Artifact Plan Use columns: command or procedure; purpose; prerequisite; environment; read-only or mutating; overhead and risk; authorization; expected artifact; execution status. Clearly distinguish suggested commands from commands actually run. ### Evidence-Gated Optimization Register Use columns: candidate; prerequisite evidence; mechanism; expected benefit; trade-offs; correctness risk; operational risk; verification method; rollout boundary; rollback trigger; decision. Set the decision to defer when prerequisite evidence is absent. ### Correctness and Performance Acceptance Matrix Use columns: criterion; baseline; target or invariant; expected observation; actual observation; evidence reference; status; owner. Do not populate actual observations unless an authorized run produced them. ### Safe Sequence and Human Decision Points List phases, entry criteria, action, owner, approval requirement, stop condition, exit evidence, and rollback or recovery action. ### Final Recommendation and Handoff State the highest-ranked unconfirmed or confirmed bottleneck, first experiment, changes to defer, safest evidence-supported path, monitoring window, rollback criteria, unresolved unknowns, and required human decisions. End with one status: blocked pending input, ready for authorized experiments, experiments executed but inconclusive, bottleneck confirmed, or optimization verified. Use the latter three only when supported by execution evidence.
Put this Prompt to work
Add the required information and prepare a version-bound task for Codex.
Opens in a new tab.
Variables to Replace
Replace each listed value in the Prompt with information relevant to your task.
- Performance symptom
- Affected workload
- Baseline evidence
- Workload model
- Code and architecture evidence
- Runtime and infrastructure
- Available profiling tools
- Test environment
- Acceptance thresholds
- Risk and authority constraints
How to Use This Prompt
Paste this prompt into Codex, replace every bracketed variable, and provide the affected code paths plus available traces, profiles, logs, query plans, benchmark results, architecture details, workload data, environment constraints, and authorization boundaries. Run the prompt in a repository or session where Codex can inspect the relevant materials; grant command execution only in an approved test environment.
Example Use Case
A reporting endpoint shows high P95 latency after a release. Codex inspects the supplied handler, ORM relationships, query samples, traces, cache telemetry, database plan evidence, and representative workload model; then it separates N+1, plan-selection, cache-churn, and serialization hypotheses into controlled experiments with safety limits and acceptance criteria.
Was this useful?