Data Analysis Expert Perplexity

Dataset and Benchmark Source Validation Brief

Produce an evidence-linked assessment of a dataset, benchmark, leaderboard result, or research metric, including provenance, methodology, freshness, comparability, limitations, claim fidelity, and fit for the intended use.

Browse more prompts
Best forSource review
ToolPerplexity
DifficultyExpert
Copied10 times
Full Prompt
Use Perplexity's web search and citation features to investigate the supplied dataset, benchmark, leaderboard result, or benchmark-based claim. Treat search results as leads, not proof: open the cited material, confirm that it supports the associated statement, and distinguish primary documentation from secondary interpretation.

## Inputs

Dataset or benchmark name: [Dataset or benchmark name]
Claim to validate: [Claim to validate]
Domain: [Domain]
Publisher or maintainer: [Publisher or maintainer]
Use case: [Use case]
Required freshness: [Required freshness]
Known concerns: [Known concerns]
Comparable benchmarks: [Comparable benchmarks]
Citation format: [Citation format]
Decision impact: [Decision impact]

Blocking inputs are the dataset or benchmark identity, the exact claim, intended use case, freshness requirement, and decision impact. If any is missing or materially ambiguous, ask focused clarification questions before issuing a use recommendation. You may still perform a bounded source-discovery pass, but label it preliminary.

The publisher, known concerns, and comparable benchmarks are useful context rather than assumed facts. If they are unknown, continue where safe and record the gap. If supplied details conflict with authoritative sources, preserve both accounts, cite the conflict, and do not resolve it without evidence. Use the requested citation format where Perplexity can support it; otherwise provide linked citations and disclose the formatting limitation.

## Tool and authority boundaries

Perplexity may search and summarize publicly accessible web sources and return citations. It may not have dependable access to private repositories, internal evaluation records, paywalled papers, deleted pages, dynamic leaderboard states, account-gated documentation, or materials blocked from indexing. Do not imply that inaccessible material was inspected. Mark each source as accessed, supplied but not independently accessed, inaccessible, or not found.

Do not publish, approve, endorse, amend, delete, license, purchase, or submit anything. Do not claim that a dataset, benchmark, model, score, citation, or public statement has been verified merely because a search result mentions it. “Verified” is permitted only when the relevant source was actually inspected, the supporting passage or artifact was identified, and the source, version, date, and evaluation conditions were reconciled with the claim. Otherwise use “partially supported,” “unverified,” “conflicting,” or “not supported.” Recommendations are advisory and require human authorization before external use.

Do not expose confidential data, credentials, personal information, unpublished evaluation material, or proprietary dataset samples. Ask for redacted excerpts or metadata when private evidence is necessary. Stop short of a definitive recommendation when identity, version, metric definition, test conditions, licensing status, or material methodology cannot be established and the decision is high impact.

## Evidence rules

1. Prioritize original benchmark papers, dataset cards, model cards, official documentation, repositories, release tags, changelogs, evaluation harnesses, leaderboard methodology, maintainer notices, licenses, and archived official pages.
2. Use independent replications, peer-reviewed critiques, audits, issue trackers, and reputable technical analyses to test—not replace—primary-source claims.
3. Label evidence as supplied fact, direct source observation, secondary report, inference, assumption, unknown, or conflict.
4. Record publication, retrieval, release, and last-update dates when available. Do not treat a page's current display date as the artifact's release date without confirmation.
5. Check that each citation resolves to the stated source and supports the nearby assertion. A citation that only mentions the subject does not validate the assertion.
6. Quote or closely paraphrase the decisive passage when practical. Do not invent scores, sample sizes, splits, confidence intervals, dates, versions, licenses, methods, or limitations.
7. For mutable leaderboards, state that the observed rank or score is a time-bounded snapshot unless an official dated record establishes otherwise.
8. Separate absence of evidence from evidence of absence.

## Investigation workflow

### 1. Normalize the validation question
Restate the exact claim as a testable proposition. Identify its subject, comparison class, metric, metric direction, value or rank, dataset or benchmark version, model or system version, evaluation date, task, split, test setup, population, geography or language, and implied scope. Record omitted qualifiers that could change its meaning.

Define the decision standard from the intended use, freshness requirement, and impact. A low-impact internal orientation may tolerate qualified secondary evidence; a customer-facing, academic, investor, regulatory, medical, financial, security, policy, or other high-impact claim requires stronger primary evidence and human review.

### 2. Establish artifact identity and provenance
Locate the canonical source and reconcile naming variants or similarly named artifacts. Determine, where available:
- original publisher, maintainer, or governing organization;
- official page, paper, repository, dataset card, benchmark card, or evaluation harness;
- release date, latest material update, version, commit, tag, DOI, or archive record;
- active, maintained, archived, superseded, deprecated, withdrawn, or unclear status;
- license, access conditions, permitted uses, redistribution limits, and material governance terms;
- lineage, source datasets, transformations, and dependencies.

Do not infer maintenance from a reachable website alone. If the original artifact has changed, distinguish the current state from the version relevant to the claim.

### 3. Inspect methodology and reproducibility
Assess the documented construction and evaluation process, including:
- collection source, sampling frame, sample size, coverage, and inclusion or exclusion rules;
- train, validation, test, hidden-test, temporal, or geographic splits;
- annotation protocol, annotator qualifications, agreement measures, adjudication, and quality controls;
- task definition, prompts or instructions, preprocessing, allowed tools, retrieval, fine-tuning, and few-shot conditions;
- metric definition, aggregation, weighting, variance, confidence intervals, significance testing, and treatment of ties;
- baseline selection and whether higher or lower values are better;
- submission policy, number of attempts, private versus public tests, and anti-gaming controls;
- evaluation code, environment, seeds, dependencies, hardware assumptions, and reproducibility artifacts;
- contamination checks, leakage controls, memorization risk, and benchmark exposure;
- known corrections, retractions, disputed labels, broken samples, or scoring changes.

For every material methodology element, report documented, partially documented, not documented, inaccessible, or not applicable. Do not convert missing documentation into a favorable finding.

### 4. Test the claim against the evidence
Decompose compound claims into atomic claims. For each one:
- identify the strongest source and exact supporting location;
- compare the claim's wording with the source's wording;
- reconcile dates, versions, model identity, metric, task, split, population, and evaluation conditions;
- determine whether the evidence is current enough;
- assess whether the claim improperly generalizes from one task, language, population, benchmark, or test setup;
- flag causal language supported only by correlation, “state of the art” language without a defined comparison set, and rank claims based on mutable or incomplete leaderboards;
- assign supported, partially supported, unverified, conflicting, or not supported;
- provide a concise reason and safer wording.

### 5. Evaluate limitations and failure modes
Address only relevant risks, but actively check for outdated or narrow data, selection and survivorship bias, demographic, geographic or language imbalance, weak labels, construct-validity problems, proxy metrics, distribution shift, contamination, leakage, overfitting, repeated submissions, leaderboard gaming, cherry-picked tasks, missing uncertainty, unfair baselines, undisclosed model assistance, nonrepresentative test conditions, poor reproducibility, maintenance uncertainty, licensing restrictions, and marketing overstatement.

For each material risk, state the evidence, likely effect on the claim or use case, severity, and mitigation or verification needed. Clearly distinguish a documented limitation from a plausible but untested concern.

### 6. Compare other evidence fairly
If comparators are supplied or discovered, first determine whether comparison is valid. Reconcile task definition, dataset version, test split, metric and direction, evaluation harness, model category, allowed resources, date, population, language, sample size, and uncertainty. Do not rank incomparable results as though they came from one controlled evaluation.

When no fair comparator is available, state which independent benchmark, replication, domain-specific evaluation, temporal holdout, external-validity study, or internal test would reduce uncertainty. Do not fabricate a comparison table from incompatible evidence.

### 7. Determine fit for purpose
Assess suitability specifically for the stated use case rather than assigning universal quality. Consider evidentiary strength, relevance, freshness, reproducibility, representativeness, licensing, consequences of error, and whether the claim can be phrased with sufficient qualification.

Choose one advisory disposition:
- Suitable to use
- Suitable with explicit caveats
- Internal context only
- Blocked pending verification
- Not suitable for this use case

State the confidence as high, moderate, low, or indeterminate and explain what evidence limits it. A disposition is not approval to publish or adopt the artifact.

### 8. Verify and reconcile before finalizing
Create acceptance checks with expected evidence, actual observation, result, and unresolved issue. At minimum verify:
- canonical artifact identity;
- source accessibility and citation support;
- relevant version and date;
- metric, task, split, and evaluation-condition match;
- methodology coverage sufficient for the decision;
- leaderboard rank or score snapshot date;
- comparator fairness;
- material limitations and licensing status;
- freshness against the stated requirement;
- safer wording consistent with the evidence.

Use pass, partial, fail, or blocked for each check. Overall acceptance requires no unresolved fail or blocked item that could materially change the proposed claim or recommendation. If sources disagree, document the competing evidence and explain what would reconcile it. Never present planned checks as completed checks.

## Required deliverable

### Validation Scope and Decision Standard
State the normalized proposition, intended use, impact, freshness threshold, blocking ambiguities, and evidence standard.

### Evidence and Provenance Ledger
Provide a table with: Evidence ID; source and URL; publisher; source type; artifact version or commit; publication or release date; last material update; access status; primary or secondary; exact fact supported; decisive passage or location; reliability notes.

### Artifact Identity and Lifecycle
Report the canonical identity, maintainer, lineage, version relevant to the claim, current maintenance state, license or usage constraints, and unresolved identity conflicts.

### Methodology and Reproducibility Matrix
Provide a table with: Methodology element; documented method; evidence ID; status; limitation; consequence for the stated use.

### Atomic Claim Check
Provide a table with: Claim ID; exact atomic claim; required qualifiers; supporting evidence; version and date match; condition match; freshness result; status; reason; safer wording.

### Limitations and Risk Register
Provide a table with: Risk or limitation; documented or plausible; evidence; severity; effect on interpretation; mitigation or further check.

### Comparator Fairness Review
Provide a table with: Comparator; common task and scope; metric compatibility; version and date alignment; evaluation-condition alignment; uncertainty available; fair comparison status; conclusion. If comparison is not defensible, explain why instead of forcing a ranking.

### Fit-for-Purpose Decision
Give the advisory disposition, confidence, rationale, allowed use with caveats, uses to avoid, and conditions that would change the decision.

### Safer Claim Wording
Provide one publication-ready qualified alternative and, when evidence is insufficient, a non-claim alternative that describes what is known without implying validation.

### Verification and Acceptance Record
Provide a table with: Check; expected evidence; actual observation; evidence ID; result; unresolved issue. Explicitly state whether the review is complete, preliminary, or blocked based on work actually performed.

### Human Review and Handoff
List materials a reviewer should inspect, unresolved conflicts, private or inaccessible evidence needed, legal or licensing questions, the person or function that should authorize consequential use, and the next verification steps. Require human review before any public-facing or high-impact claim is published or relied upon.

Finish with a short conclusion naming the strongest evidence, weakest material evidence, primary limitation, recommended disposition, confidence, and the single most important unresolved check.

Put this Prompt to work

Add the required information and run this Prompt with your selected AI provider.

Opens in a new tab.

Variables to Replace

Replace each listed value in the Prompt with information relevant to your task.

  • Dataset or benchmark name
  • Claim to validate
  • Domain
  • Publisher or maintainer
  • Use case
  • Required freshness
  • Known concerns
  • Comparable benchmarks
  • Citation format
  • Decision impact

How to Use This Prompt

Paste the prompt into Perplexity, replace every bracketed variable, and provide the exact claim plus relevant source materials such as official dataset pages, papers, dataset or benchmark cards, repository links, release notes, evaluation harnesses, leaderboard snapshots, licenses, and internal redacted evidence. Run the prompt, inspect the cited passages and unresolved items, then obtain human review before using any public-facing or high-impact conclusion.

Example Use Case

A data team plans to cite a model's “number one” benchmark ranking in a customer report. The brief checks the dated leaderboard snapshot, benchmark and model versions, metric and test conditions, methodology, contamination risk, comparator fairness, and source support before recommending qualified wording or blocking the claim pending evidence.

Was this useful?

Build stronger AI systems

Use Amo.ng prompts as reusable building blocks, then go deeper with RichlyAI.

Related Prompts

Browse all
Data Analysis Expert Codex

Data Lineage Break Investigation

Reconstruct where a data product diverged from authoritative lineage, bound affected outputs and decisions, and define safe repair and reprocessing.

Updated Aug 25, 2026

View prompt Verified ✓ 201 views · 18 copies