Source version 1.0.0
Published
Initial: Initial published snapshot.
Published version comparison
1.0.0 → 2.0.0
1.0.0Published
Initial: Initial published snapshot.
2.0.0Published
Major: Replace the legacy Dataset and Benchmark Source Validation Brief template with a domain-specific input, evidence, authority, safety, workflow, output, and verification contract.
Dataset and Benchmark Source Validation Brief
Dataset and Benchmark Source Validation Brief
Validate datasets, benchmarks, leaderboard claims, and research metrics with source checks, methodology review, limitations, freshness, and usage risks.
Produce an evidence-linked assessment of a dataset, benchmark, leaderboard result, or research metric, including provenance, methodology, freshness, comparability, limitations, claim fidelity, and fit for the intended use.
Checking whether a dataset, benchmark, leaderboard, or research claim is credible, current, well-sourced, and suitable for analysis or public claims.
Checking whether a dataset, benchmark, leaderboard result, or related research claim is sufficiently credible, current, comparable, and well-supported for a specific decision or publication.
Dataset Source Validation Benchmark Methodology Review Leaderboard Claim Checking Research Evidence Review Data Limitation Analysis Fit-for-Purpose Assessment
Validating a dataset's provenance, version, license, and fitness for a defined analysis Reviewing benchmark methodology, reproducibility, leakage, and contamination controls Checking dated leaderboard rank or score claims against official evaluation conditions Testing whether benchmark comparisons use compatible tasks, metrics, splits, and versions Preparing qualified evidence language for customer-facing, academic, or executive claims Identifying verification gaps before a high-impact dataset or benchmark decision
Dataset or benchmark name Claim to validate Domain Publisher or maintainer Use case Required freshness Known concerns Comparable benchmarks Citation format Decision impact
Dataset or benchmark name Claim to validate Domain Publisher or maintainer Use case Required freshness Known concerns Comparable benchmarks Citation format Decision impact
Paste this prompt into Perplexity with the dataset or benchmark name, claim to validate, domain, publisher, intended use case, freshness requirement, known concerns, comparable benchmarks, citation format, and decision impact. Use the output to decide whether the dataset, benchmark, or leaderboard claim is reliable enough to cite, compare, or use in decision-making.
Paste the prompt into Perplexity, replace every bracketed variable, and provide the exact claim plus relevant source materials such as official dataset pages, papers, dataset or benchmark cards, repository links, release notes, evaluation harnesses, leaderboard snapshots, licenses, and internal redacted evidence. Run the prompt, inspect the cited passages and unresolved items, then obtain human review before using any public-facing or high-impact conclusion.
A data team wants to cite an AI benchmark in a customer-facing report. The prompt checks the original source, methodology, benchmark date, leaderboard claim, limitations, comparable evidence, and safer wording before the claim is published.
A data team plans to cite a model's “number one” benchmark ranking in a customer report. The brief checks the dated leaderboard snapshot, benchmark and model versions, metric and test conditions, methodology, contamination risk, comparator fairness, and source support before recommending qualified wording or blocking the claim pending evidence.
Expert
Expert
Perplexity
Perplexity
source review
source review
dataset-validation benchmark-review perplexity source-review methodology-review leaderboard-claims research-integrity data-quality evidence-check limitation-analysis
dataset-validation benchmark-methodology leaderboard-claims source-provenance evidence-review perplexity reproducibility benchmark-contamination fit-for-purpose research-integrity
Dataset and Benchmark Source Validation Brief Prompt
Dataset and Benchmark Source Validation Brief
Validate datasets, benchmarks, leaderboard claims, and research metrics with source checks, methodology review, limitations, freshness, and usage risks.
Validate dataset, benchmark, and leaderboard claims through provenance, methodology, freshness, comparability, limitations, and citation checks.
Removed Added Unchanged context
You are an expert research methods analyst specializing in dataset validation, benchmark review, leaderboard claim assessment, source provenance, methodology analysis, data quality, research integrity, and fit-for-purpose evidence review. Use Perplexity's web search and citation features to investigate the supplied dataset, benchmark, leaderboard result, or benchmark-based claim. Treat search results as leads, not proof: open the cited material, confirm that it supports the associated statement, and distinguish primary documentation from secondary interpretation. Your task is to assess whether a dataset, benchmark, leaderboard, or benchmark-based claim is credible, current, well-sourced, methodologically sound, and appropriate for the intended decision or public claim. ## Inputs Context: Dataset or benchmark name: [Dataset or benchmark name] Claim to validate: [Claim to validate] Domain: [Domain] Publisher or maintainer: [Publisher or maintainer] Use case: [Use case] Required freshness: [Required freshness] Known concerns: [Known concerns] Comparable benchmarks: [Comparable benchmarks] Citation format: [Citation format] Decision impact: [Decision impact] Important constraints: * Use source-backed reasoning. * Prioritize primary sources such as official dataset pages, benchmark papers, documentation, methodology notes, repository pages, release notes, and maintainer announcements. * Do not invent dataset details, benchmark scores, citations, publication dates, sample sizes, methodology claims, licensing terms, or limitations. * Separate confirmed information from assumptions. * Clearly distinguish official sources from commentary, summaries, blog posts, marketing claims, and secondary interpretations. * Do not recommend using a dataset or benchmark without naming its limitations and fit-for-purpose concerns. * Check whether the benchmark or dataset is current enough for the stated use case. * Check whether the benchmark claim is being overstated beyond what the source supports. * Include human review for public-facing, investor-facing, legal, regulatory, academic, medical, financial, technical, security, or high-impact claims. * If source information is missing or unclear, mark it as “Needs verification.” * If the available evidence is insufficient, say so clearly. Task: 1. Summarize the validation objective. Explain: * Dataset or benchmark being reviewed * Claim being validated * Domain * Intended use case * Decision impact * Required freshness * Main source question to answer 2. Review source provenance. Identify: Blocking inputs are the dataset or benchmark identity, the exact claim, intended use case, freshness requirement, and decision impact. If any is missing or materially ambiguous, ask focused clarification questions before issuing a use recommendation. You may still perform a bounded source-discovery pass, but label it preliminary. * Original publisher or maintainer * Official source URL or citation * Publication or release date * Latest update date, if available * Version number, if available * Repository or documentation location * Whether the source appears active, archived, deprecated, or unclear * Whether the cited source is primary, secondary, or commentary The publisher, known concerns, and comparable benchmarks are useful context rather than assumed facts. If they are unknown, continue where safe and record the gap. If supplied details conflict with authoritative sources, preserve both accounts, cite the conflict, and do not resolve it without evidence. Use the requested citation format where Perplexity can support it; otherwise provide linked citations and disclose the formatting limitation. Create a table with: ## Tool and authority boundaries * Source * Source type * Date * What it supports * Reliability level * Notes or concerns Perplexity may search and summarize publicly accessible web sources and return citations. It may not have dependable access to private repositories, internal evaluation records, paywalled papers, deleted pages, dynamic leaderboard states, account-gated documentation, or materials blocked from indexing. Do not imply that inaccessible material was inspected. Mark each source as accessed, supplied but not independently accessed, inaccessible, or not found. 3. Review methodology. Assess: Do not publish, approve, endorse, amend, delete, license, purchase, or submit anything. Do not claim that a dataset, benchmark, model, score, citation, or public statement has been verified merely because a search result mentions it. “Verified” is permitted only when the relevant source was actually inspected, the supporting passage or artifact was identified, and the source, version, date, and evaluation conditions were reconciled with the claim. Otherwise use “partially supported,” “unverified,” “conflicting,” or “not supported.” Recommendations are advisory and require human authorization before external use. * How the dataset or benchmark was created * Data collection method * Sample size or scope, if available * Evaluation method * Scoring method * Task definition * Inclusion and exclusion criteria * Annotation or labeling process, if relevant * Validation process * Reproducibility details * Known methodological weaknesses Do not expose confidential data, credentials, personal information, unpublished evaluation material, or proprietary dataset samples. Ask for redacted excerpts or metadata when private evidence is necessary. Stop short of a definitive recommendation when identity, version, metric definition, test conditions, licensing status, or material methodology cannot be established and the decision is high impact. If methodology details are missing, mark them as “Needs verification.” ## Evidence rules 4. Check benchmark or leaderboard claims. For each claim, determine: 1. Prioritize original benchmark papers, dataset cards, model cards, official documentation, repositories, release tags, changelogs, evaluation harnesses, leaderboard methodology, maintainer notices, licenses, and archived official pages. 2. Use independent replications, peer-reviewed critiques, audits, issue trackers, and reputable technical analyses to test—not replace—primary-source claims. 3. Label evidence as supplied fact, direct source observation, secondary report, inference, assumption, unknown, or conflict. 4. Record publication, retrieval, release, and last-update dates when available. Do not treat a page's current display date as the artifact's release date without confirmation. 5. Check that each citation resolves to the stated source and supports the nearby assertion. A citation that only mentions the subject does not validate the assertion. 6. Quote or closely paraphrase the decisive passage when practical. Do not invent scores, sample sizes, splits, confidence intervals, dates, versions, licenses, methods, or limitations. 7. For mutable leaderboards, state that the observed rank or score is a time-bounded snapshot unless an official dated record establishes otherwise. 8. Separate absence of evidence from evidence of absence. * Exact claim being made * Source supporting the claim * Whether the claim matches the source * Whether the claim is current * Whether the claim depends on a specific version, date, model, task, metric, or test setup * Whether the claim is being overstated * Safer wording for the claim ## Investigation workflow 5. Identify limitations and risks. Review possible issues such as: ### 1. Normalize the validation question Restate the exact claim as a testable proposition. Identify its subject, comparison class, metric, metric direction, value or rank, dataset or benchmark version, model or system version, evaluation date, task, split, test setup, population, geography or language, and implied scope. Record omitted qualifiers that could change its meaning. * Outdated data * Small or narrow sample * Domain mismatch * Selection bias * Geographic bias * Language bias * Demographic bias * Labeling quality issues * Benchmark contamination * Data leakage * Overfitting to benchmark tasks * Non-representative test conditions * Licensing or usage restrictions * Unclear maintenance * Missing documentation * Poor reproducibility * Leaderboard gaming * Marketing overclaiming Define the decision standard from the intended use, freshness requirement, and impact. A low-impact internal orientation may tolerate qualified secondary evidence; a customer-facing, academic, investor, regulatory, medical, financial, security, policy, or other high-impact claim requires stronger primary evidence and human review. 6. Compare with other evidence. If comparable benchmarks or datasets are provided, compare: ### 2. Establish artifact identity and provenance Locate the canonical source and reconcile naming variants or similarly named artifacts. Determine, where available: - original publisher, maintainer, or governing organization; - official page, paper, repository, dataset card, benchmark card, or evaluation harness; - release date, latest material update, version, commit, tag, DOI, or archive record; - active, maintained, archived, superseded, deprecated, withdrawn, or unclear status; - license, access conditions, permitted uses, redistribution limits, and material governance terms; - lineage, source datasets, transformations, and dependencies. * Scope * Methodology * Freshness * Credibility * Known limitations * Use-case fit * Whether the comparison is fair Do not infer maintenance from a reachable website alone. If the original artifact has changed, distinguish the current state from the version relevant to the claim. If no comparable evidence is provided, suggest what type of comparison should be checked before relying on the claim. ### 3. Inspect methodology and reproducibility Assess the documented construction and evaluation process, including: - collection source, sampling frame, sample size, coverage, and inclusion or exclusion rules; - train, validation, test, hidden-test, temporal, or geographic splits; - annotation protocol, annotator qualifications, agreement measures, adjudication, and quality controls; - task definition, prompts or instructions, preprocessing, allowed tools, retrieval, fine-tuning, and few-shot conditions; - metric definition, aggregation, weighting, variance, confidence intervals, significance testing, and treatment of ties; - baseline selection and whether higher or lower values are better; - submission policy, number of attempts, private versus public tests, and anti-gaming controls; - evaluation code, environment, seeds, dependencies, hardware assumptions, and reproducibility artifacts; - contamination checks, leakage controls, memorization risk, and benchmark exposure; - known corrections, retractions, disputed labels, broken samples, or scoring changes. 7. Assess fit for the intended use case. Evaluate whether the dataset or benchmark is suitable for: For every material methodology element, report documented, partially documented, not documented, inaccessible, or not applicable. Do not convert missing documentation into a favorable finding. * Internal research * Public article or report * Academic citation * Product comparison * Model evaluation * Customer-facing claim * Investor or executive presentation * Policy, compliance, or high-impact decision ### 4. Test the claim against the evidence Decompose compound claims into atomic claims. For each one: - identify the strongest source and exact supporting location; - compare the claim's wording with the source's wording; - reconcile dates, versions, model identity, metric, task, split, population, and evaluation conditions; - determine whether the evidence is current enough; - assess whether the claim improperly generalizes from one task, language, population, benchmark, or test setup; - flag causal language supported only by correlation, “state of the art” language without a defined comparison set, and rank claims based on mutable or incomplete leaderboards; - assign supported, partially supported, unverified, conflicting, or not supported; - provide a concise reason and safer wording. Explain what level of confidence is justified. ### 5. Evaluate limitations and failure modes Address only relevant risks, but actively check for outdated or narrow data, selection and survivorship bias, demographic, geographic or language imbalance, weak labels, construct-validity problems, proxy metrics, distribution shift, contamination, leakage, overfitting, repeated submissions, leaderboard gaming, cherry-picked tasks, missing uncertainty, unfair baselines, undisclosed model assistance, nonrepresentative test conditions, poor reproducibility, maintenance uncertainty, licensing restrictions, and marketing overstatement. 8. Create a use recommendation. Classify the dataset, benchmark, or claim as one of: For each material risk, state the evidence, likely effect on the claim or use case, severity, and mitigation or verification needed. Clearly distinguish a documented limitation from a plausible but untested concern. * Suitable to use * Suitable with caveats * Use only for internal context * Do not use without further verification * Not suitable for this use case ### 6. Compare other evidence fairly If comparators are supplied or discovered, first determine whether comparison is valid. Reconcile task definition, dataset version, test split, metric and direction, evaluation harness, model category, allowed resources, date, population, language, sample size, and uncertainty. Do not rank incomparable results as though they came from one controlled evaluation. Explain the reason clearly. When no fair comparator is available, state which independent benchmark, replication, domain-specific evaluation, temporal holdout, external-validity study, or internal test would reduce uncertainty. Do not fabricate a comparison table from incompatible evidence. 9. Provide safer claim wording. Rewrite the original claim into a more accurate version that reflects: ### 7. Determine fit for purpose Assess suitability specifically for the stated use case rather than assigning universal quality. Consider evidentiary strength, relevance, freshness, reproducibility, representativeness, licensing, consequences of error, and whether the claim can be phrased with sufficient qualification. * Source limits * Date or version * Methodology constraints * Scope * Uncertainty * Caveats Choose one advisory disposition: - Suitable to use - Suitable with explicit caveats - Internal context only - Blocked pending verification - Not suitable for this use case 10. Provide final recommendations. Summarize: State the confidence as high, moderate, low, or indeterminate and explain what evidence limits it. A disposition is not approval to publish or adopt the artifact. * Best available source * Strongest supporting evidence * Weakest evidence * Main limitations * Freshness concerns * Fit-for-purpose concerns * Human review needed * Next verification steps before citing or using the claim ### 8. Verify and reconcile before finalizing Create acceptance checks with expected evidence, actual observation, result, and unresolved issue. At minimum verify: - canonical artifact identity; - source accessibility and citation support; - relevant version and date; - metric, task, split, and evaluation-condition match; - methodology coverage sufficient for the decision; - leaderboard rank or score snapshot date; - comparator fairness; - material limitations and licensing status; - freshness against the stated requirement; - safer wording consistent with the evidence. Output format: Use pass, partial, fail, or blocked for each check. Overall acceptance requires no unresolved fail or blocked item that could materially change the proposed claim or recommendation. If sources disagree, document the competing evidence and explain what would reconcile it. Never present planned checks as completed checks. ## Validation Objective ## Required deliverable ## Source Provenance ### Validation Scope and Decision Standard State the normalized proposition, intended use, impact, freshness threshold, blocking ambiguities, and evidence standard. ## Methodology Review ### Evidence and Provenance Ledger Provide a table with: Evidence ID; source and URL; publisher; source type; artifact version or commit; publication or release date; last material update; access status; primary or secondary; exact fact supported; decisive passage or location; reliability notes. ## Benchmark or Leaderboard Claim Check ### Artifact Identity and Lifecycle Report the canonical identity, maintainer, lineage, version relevant to the claim, current maintenance state, license or usage constraints, and unresolved identity conflicts. ## Limitations and Risks ### Methodology and Reproducibility Matrix Provide a table with: Methodology element; documented method; evidence ID; status; limitation; consequence for the stated use. ## Comparable Evidence ### Atomic Claim Check Provide a table with: Claim ID; exact atomic claim; required qualifiers; supporting evidence; version and date match; condition match; freshness result; status; reason; safer wording. ## Fit-for-Purpose Assessment ### Limitations and Risk Register Provide a table with: Risk or limitation; documented or plausible; evidence; severity; effect on interpretation; mitigation or further check. ## Use Recommendation ### Comparator Fairness Review Provide a table with: Comparator; common task and scope; metric compatibility; version and date alignment; evaluation-condition alignment; uncertainty available; fair comparison status; conclusion. If comparison is not defensible, explain why instead of forcing a ranking. ## Safer Claim Wording ### Fit-for-Purpose Decision Give the advisory disposition, confidence, rationale, allowed use with caveats, uses to avoid, and conditions that would change the decision. ## Final Recommendations ### Safer Claim Wording Provide one publication-ready qualified alternative and, when evidence is insufficient, a non-claim alternative that describes what is known without implying validation. Verification: Before finalizing, check that: ### Verification and Acceptance Record Provide a table with: Check; expected evidence; actual observation; evidence ID; result; unresolved issue. Explicitly state whether the review is complete, preliminary, or blocked based on work actually performed. * Every factual claim is tied to a source. * Source dates and versions are included where available. * Primary sources are prioritized over commentary. * Methodology limitations are clearly stated. * Dataset or benchmark freshness is assessed. * Benchmark claims are not overstated. * Fit-for-purpose concerns are named. * Human review is recommended for high-impact or public-facing claims. * Missing information is marked as “Needs verification.” * The recommendation is cautious when evidence is incomplete. ### Human Review and Handoff List materials a reviewer should inspect, unresolved conflicts, private or inaccessible evidence needed, legal or licensing questions, the person or function that should authorize consequential use, and the next verification steps. Require human review before any public-facing or high-impact claim is published or relied upon. Begin the dataset and benchmark source validation brief now. Finish with a short conclusion naming the strongest evidence, weakest material evidence, primary limitation, recommended disposition, confidence, and the single most important unresolved check.