EvidenceForge is a skill/workflow project, not a replacement for formal methods guidance. This file records the methodological sources that shape the current skills.
- Xu, Yiqing and Leo Yang Yang. Scaling Reproducibility: An AI-Assisted Workflow for Large-Scale Replication and Reanalysis.
EvidenceForge borrows the workflow architecture: skill contracts, explicit intermediate artifacts, AI orchestration separated from deterministic analysis, and human judgment kept visible. The meta-analysis and evidence-synthesis direction itself comes from the author's research experience and from the broader spread of meta-analysis beyond medicine into ecology, environment, life science, policy, and social science.
Source: PRISMA 2020
Use for:
- transparent reporting;
- checklist-driven manuscripts;
- abstract checklist;
- flow diagrams for original and updated reviews;
- making search, screening, inclusion, and synthesis decisions visible.
EvidenceForge implication: every review report should be able to produce a PRISMA-style account of search results, screening stages, exclusions, and included studies.
Source: Cochrane Handbook for Systematic Reviews of Interventions
Use for:
- planning a review;
- defining scope and inclusion criteria;
- searching and selecting studies;
- data collection;
- effect measures;
- risk of bias;
- meta-analysis;
- network meta-analysis;
- non-meta-analytic synthesis;
- missing-result bias;
- Summary of Findings and GRADE;
- interpreting results.
EvidenceForge implication: statistical synthesis should be chosen after the review question, eligibility, data extraction, risk-of-bias, and effect-measure decisions are clear.
Source: JBI Manual for Evidence Synthesis
Use for:
- broad evidence-synthesis methodology;
- systematic reviews of effectiveness, qualitative evidence, textual evidence, economic evidence, etiology/risk;
- mixed-methods reviews;
- umbrella reviews;
- scoping reviews;
- inclusive approaches across diverse questions and designs.
EvidenceForge implication: not every evidence project is a conventional intervention meta-analysis. The review type should match the question.
Source: CEE Summary of Standards
Use for:
- environmental systematic reviews and systematic maps;
- a priori protocols;
- PICO/PECO-style question elements;
- grey literature and supplementary searching;
- search comprehensiveness tests;
- dual screening and consistency checks;
- structured data extraction;
- critical appraisal;
- deciding when meta-analysis is inappropriate.
EvidenceForge implication: environmental and ecological reviews need explicit search transparency, geography/context fields, effect modifiers, study validity appraisal, and caution against vote-counting.
Source family: R pls package workflows and VIP-based environmental indicator modeling.
Use for:
- NDVI, vegetation, soil, climate, hydrology, biodiversity, and ecosystem indicator models;
- partial least squares regression with correlated predictors and modest samples;
- component-selection and cross-validation audits;
- VIP ranking and coefficient-direction interpretation;
- separating predictive variable contribution from causal or mechanistic claims.
EvidenceForge implication: PLS/VIP studies should be audited as predictive or exploratory indicator models unless a separate causal design is present. Agents should require scaling, component-selection logic, validation design, residual checks, VIP tables, and direction of association before summarizing ecological meaning.
Source: Ren, Liu, and Lin (2026), Threshold-oriented ecosystem-service relationships from interpretable machine learning to spatial management: A case from Shandong Province, China, Ecological Indicators.
Use for:
- ecosystem-service relationship mapping;
- synergy/trade-off classification across pairwise service combinations;
- GWR and spatial heterogeneity audits;
- XGBoost, GBDT, GAM, and MLR model comparison;
- nonlinear driver-response interpretation;
- partial-response curve superposition;
- optimal interval and threshold identification;
- translating model patterns into zoning, agricultural upgrading, urban-rural interface control, and biodiversity-friendly conservation guidance.
EvidenceForge implication: interpretable ML can make ecosystem-service analysis more operational, but threshold intervals should be treated as probability-oriented management heuristics unless supported by causal, temporal, experimental, or mechanistic evidence. Agents should audit service-map uncertainty, spatial leakage, model validation, threshold stability, and local feasibility before turning model-derived intervals into planning recommendations.
Source: Liu et al. (2024), Air quality improvements can strengthen China's food security, Nature Food.
Use for:
- ozone, PM2.5, aerosol, SIF, crop productivity, and yield-response modeling;
- satellite-based observations with flexible statistical forms;
- counterfactual air-quality target scenarios;
- translating crop-yield changes into crop calories and food-security implications;
- code/data availability auditing for policy co-benefit models.
EvidenceForge implication: air-quality food-security studies should be audited as a chain of linked assumptions. Agents should separately check pollutant exposure, crop productivity proxy, yield conversion, calorie conversion, and self-sufficiency or policy claims.
Source seed: Qi et al. (2026), Intensifying Aridity Undermines the Role of Soil Biodiversity in Supporting Ecosystem Stability, Global Change Biology, DOI https://doi.org/10.1111/gcb.70903 from user-provided screenshot.
Use for:
- biodiversity-stability relationship audits;
- soil biodiversity and ecosystem stability evidence;
- aridity or climate-stress moderation;
- scale, confounding, and mechanism checks;
- distinguishing observational, experimental, and causal claims.
EvidenceForge implication: biodiversity-stability claims need the stability metric, biodiversity dimension, stress gradient, and temporal/spatial scale to be explicit. When aridity is a moderator, agents should avoid summarizing the paper as a simple biodiversity benefit unless the stress-dependent relationship is preserved.
Source: Li et al. (2026), The underappreciated importance of small wetlands in global methane emissions, Nature Climate Change.
Use for:
- small-wetland and small-water-body emission synthesis;
- fine-resolution remote sensing and minimum-mapping-unit audits;
- wetland methane budget corrections;
- size-class accounting for 0.001-1 km2 and <0.1 km2 features;
- trend analysis for 2003-2022;
- data/code availability checks for geospatial emission studies.
EvidenceForge implication: scale is part of the evidence claim. Agents should audit what coarse maps omit, how size classes are defined, whether double-counting occurs across wetlands/lakes/rivers/reservoirs, and whether methane-only results are being overextended to wetland management.
Source: Wang, Ran, Li, et al. (2026), Near-surface ground ice map of the Northern Hemisphere, Science Bulletin.
Use for:
- permafrost and near-surface ground-ice map-product audits;
- borehole observation inventories and depth-convention checks;
- 1 km gridded cryosphere mapping;
- predictor stacks that combine substrate, hydrology, topography, surficial geology, paleoclimate, modern climate, remote sensing, and vegetation;
- ensemble machine-learning and Bayesian model-averaging workflows;
- spatial autocorrelation, extrapolation, and validation checks;
- prediction intervals and map-uncertainty interpretation;
- public raster/data DOI and versioning requirements.
EvidenceForge implication: environmental evidence products are not just papers; they can be maps with observation, prediction, and interpretation layers. Agents should keep those layers separate and should not translate predicted ice content into climate, hydrology, ecosystem, or infrastructure claims without uncertainty, spatial validation, and local-context checks.
Source: Mogollón et al. (2026), Broad bidirectional effects of global food production on the environment, Nature Reviews Earth & Environment.
Use for:
- food-system environmental pressure mapping;
- bidirectional links between production impacts and environmental feedbacks;
- separating crop, livestock, and aquatic/blue-food receptors;
- production-based and consumption-based accounting;
- trade displacement and embodied environmental impacts;
- supply-side, demand-side, adaptation, and circularity strategy maps.
EvidenceForge implication: food-system evidence synthesis should not stop at footprint accounting. Agents should map both directions: food production pressures on the environment and environmental degradation feedbacks on future production and food security.
Source: Uen and Rodriguez (2026), Geospatial analytics and machine learning for forecasting county-level food waste in U.S. Retail markets, Resources, Conservation and Recycling.
Use for:
- county-level food-waste forecasting;
- geospatial machine-learning audit;
- feature-importance interpretation for demographic, retail, expenditure, income, and social-program predictors;
- spatial validation and high-waste outlier checks;
- translating predictions into logistics, facility siting, waste-to-bioproduct systems, and circular bioeconomy planning.
EvidenceForge implication: food-waste ML studies need a planning-translation layer. Agents should distinguish generated waste from collectible and usable feedstock, and should not treat county-level model performance as enough for routing, facility siting, or policy claims without spatial and logistical validation.
Source: Wang et al. (2025), Integrating meteorological and breeding data to predict maize yields using machine learning algorithms, Frontiers in Plant Science.
Use for:
- crop-yield prediction with meteorological and breeding data;
- genotype-by-environment data integration;
- BLUP breeding values as model inputs;
- comparing Random Forest, XGBoost, SVR, and GPR;
- validation split and leakage auditing;
- feature-importance and partial-dependence interpretation;
- distinguishing operational prediction from causal agronomic explanation.
EvidenceForge implication: agricultural ML papers need model-audit templates that track site, year, genotype, weather windows, validation groups, and deployment claims. Random train/test accuracy is not enough when the claim involves new sites, years, genotypes, or future climates.
Source: Liang and Schlesinger (2026), Potassium fertilization enhances both cereal yield and soil organic carbon: a meta-analysis, Nature Communications.
Use for:
- agricultural nutrient-management meta-analysis;
- crop yield and soil organic carbon as separate but policy-linked endpoints;
- response-ratio and percent-change interpretation;
- moderator extraction for climate, baseline soil status, and experimental duration;
- long-term evidence gates for soil-carbon claims;
- open data/code audit practices.
EvidenceForge implication: environmental meta-analysis often needs dual-outcome interpretation. A review may show a positive yield response and a smaller or slower soil-carbon response; the skill should keep time horizon, soil depth, baseline nutrient limitation, and study dependence visible.
Source: Ba et al. (2026), Full Manure Recycling Risks an 18% Rise in China's Cropland N2O Emissions Without Improved Management, Global Change Biology.
Use for:
- literature-derived environmental databases;
- measurement-level extraction from many field studies;
- machine-learning prediction with validation;
- gridded spatial extrapolation;
- future policy scenario design;
- separating policy coverage from implementation quality.
EvidenceForge implication: environmental evidence synthesis sometimes needs a scenario-model audit rather than a pooled effect estimate. Agents should track database construction, model validation, spatial assumptions, scenario levers, uncertainty, and policy trade-offs.
Source: Bayer, Lautenbach, and Arneth (2023), Benefits and trade-offs of optimizing global land use for food, water, and carbon, PNAS.
Use for:
- multiobjective ecosystem-service optimization;
- Pareto-frontier / production-possibility frontier interpretation;
- dynamic vegetation model outputs as optimization inputs;
- food-water-carbon trade-off mapping;
- spatial priority maps and solution-agreement audits;
- separating theoretical biogeophysical option space from feasible policy pathways.
EvidenceForge implication: some environmental evidence projects should ask what the modeled option space is, not just what one scenario predicts. Agents should audit objective definitions, units, baselines, constraints, optimization algorithms, transition costs, omitted social constraints, and local-vs-global interpretation.
Source: AMSTAR 2
Use for:
- appraising systematic reviews of randomized and non-randomized healthcare intervention studies;
- identifying critical and non-critical weaknesses;
- confidence categories rather than a simple numerical score.
EvidenceForge implication: umbrella reviews should not treat all included reviews as equal. Review quality affects interpretation.
Source: ROBIS
Use for:
- assessing risk of bias in systematic reviews themselves;
- review-of-reviews and umbrella-review contexts.
EvidenceForge implication: second-order synthesis must evaluate bias in the review process, not only bias in primary studies.
Source: Raveloaritiana and Wanger (2026), Long-term agricultural diversification increases financial profitability, biodiversity, and ecosystem services: a second-order meta-analysis, Nature Communications.
Use for:
- second-order synthesis of existing meta-analyses;
- temporal meta-regression with duration as moderator;
- review-level effect-size extraction and weighting;
- yield vs ecosystem-service trade-off categories;
- long-term evidence-gap mapping;
- separating review-level evidence patterns from primary-study causal claims.
EvidenceForge implication: second-order meta-analysis needs its own audit layer. Agents should track review-level units, primary-study overlap, quality, metric compatibility, duration-model choices, and whether policy claims overgeneralize from highly aggregated evidence.
Source: GRADE Handbook via Cochrane
Use for:
- rating certainty or confidence in evidence;
- Summary of Findings logic;
- distinguishing effect estimates from certainty of evidence.
EvidenceForge implication: a statistically significant pooled effect does not automatically mean high-certainty evidence.
Source: ASReview AI
Use for:
- active-learning screening;
- researcher/human-in-the-loop review;
- transparent logging of human and AI decisions;
- reproducible screening workflows.
EvidenceForge implication: ML assistance should prioritize and accelerate screening while preserving audit trails and human decision authority.
Source: Rayyan
Use for:
- systematic-review management;
- AI-assisted screening workflows;
- transparent review process support.
EvidenceForge implication: platform-specific workflows should still export decisions, exclusion reasons, and review logs.
- Borenstein, Hedges, Higgins, and Rothstein, Introduction to Meta-Analysis.
- Cooper, Hedges, and Valentine, The Handbook of Research Synthesis and Meta-Analysis.
- Hedges and Olkin, Statistical Methods for Meta-Analysis.
- Harrer, Cuijpers, Furukawa, and Ebert, Doing Meta-Analysis with R.
- Koricheva, Gurevitch, and Mengersen, Handbook of Meta-analysis in Ecology and Evolution.
EvidenceForge implication: future deterministic scripts should likely begin with R workflows around metafor, meta, clubSandwich, robumeta, and related packages, while preserving the skill layer as the human-readable contract.