{"publication_id":"5c993ba1-5ebb-4a12-b4dc-a4fe2418a927","traces":[{"claim_id":"claim_1","claim":"retrieval augmented generation: Bounded signal: retrieval augmented generation is only a source-level context map; the selected receipts do not establish one pooled effect. Context-only rows are adjacent scope, not effect support; no pooled causal, policy-prescriptive, or market-generalized claim is made.","citation_support":[],"candidate_sources":[{"study":"A Retrieval-Augmented Generation Framework for Traditional Chinese Medicine Herb Recommendation Using Symptom-Focused and Ingredient-Based Embeddings","year":2026,"doi":"10.65205/jcct.2026.e3516","url":"https://doi.org/10.65205/jcct.2026.e3516","population":"rag accuracy tasks","intervention_or_exposure":"Retrieval-Augmented Generation Framework","comparator":"not extracted","endpoint":"not extracted","effect":"not extracted","risk_of_bias":"not appraised in public sidecar","directness":"primary","excerpt":"The baseline LLM demonstrated strong performance across multiple metrics, including accuracy (0.1900) and NDCG@5 (0.1475), reflecting substantial pre-trained medical knowledge.","source_id":"source_1","support_kind":"candidate_source_row"},{"study":"Evaluating Retrieval-Augmented Generation Variants for Natural Language-Based SQL and API Call Generation","year":2026,"doi":"10.48550/arxiv.2602.07086","url":"https://doi.org/10.48550/arxiv.2602.07086","population":"combined","intervention_or_exposure":"RAG","comparator":"not extracted","endpoint":"not extracted","effect":"not extracted","risk_of_bias":"not appraised in public sidecar","directness":"primary","excerpt":"Critically, CoRAG proves most robust in hybrid documentation settings, achieving statistically significant improvements in the combined task (10.29% exact match vs. 7.45% for standard RAG), driven primarily by superior SQL generation performance (15.32% vs. 11.56%).","source_id":"source_2","support_kind":"candidate_source_row"},{"study":"A retrieval-augmented generation large language model framework for accurate dementia identification from electronic health records","year":2026,"doi":"10.64898/2026.01.24.26344477","url":"https://doi.org/10.64898/2026.01.24.26344477","population":"rag F1 tasks","intervention_or_exposure":"RAG","comparator":"not extracted","endpoint":"not extracted","effect":"not extracted","risk_of_bias":"not appraised in public sidecar","directness":"primary","excerpt":"ResultsThe RAG-based classifier achieved the highest performance (F1=0.933, sensitivity=91.1%, PPV=95.5%) compared to rule-based (F1=0.823, sensitivity=81.1%, PPV=83.5%) and keyword-filtered LLM (F1=0.903, sensitivity=91.7%, PPV=88.6%).","source_id":"source_3","support_kind":"candidate_source_row"},{"study":"Integrating Dense, Sparse, and Graph-Based Approaches in Financial Data Analysis for a Retrieval-Augmented Generation Framework","year":2026,"doi":"10.1109/acdsa67686.2026.11467963","url":"https://doi.org/10.1109/acdsa67686.2026.11467963","population":"rag recall tasks","intervention_or_exposure":"Integrating Dense, Sparse, and Graph-Based Approaches","comparator":"not extracted","endpoint":"not extracted","effect":"not extracted","risk_of_bias":"not appraised in public sidecar","directness":"primary","excerpt":"Results show that integrating a graph-based retriever improved context recall by 63%, answer correctness by 31%, and overall performance by 12% compared to flattened text retrieval.","source_id":"source_4","support_kind":"candidate_source_row"},{"study":"Improving Retrieval-Augmented Generation Performance Using the MAF-RAG Architecture, EVR–VOR Vector Retrieval, and Multi-Agent Fallback Reasoning","year":2026,"doi":"10.30871/jaic.v10i1.11738","url":"https://doi.org/10.30871/jaic.v10i1.11738","population":"rag F1 tasks","intervention_or_exposure":"RAG","comparator":"not extracted","endpoint":"not extracted","effect":"not extracted","risk_of_bias":"not appraised in public sidecar","directness":"primary","excerpt":"The results show that the proposed MAF-RAG significantly outperforms the baseline system, achieving a mean F1-score of 0.556, an improvement of 18.8% over the Enhanced Baseline (mean F1-score = 0.469) and a 70.0% improvement over the Legacy Baseline (mean F1-score = 0.327).","source_id":"source_5","support_kind":"candidate_source_row"}]},{"claim_id":"claim_2","claim":"Does retrieval augmented generation show a consistent direction-bearing association in the selected source bundle, and where do null/mixed or context-only receipts bound the claim?","citation_support":[],"candidate_sources":[{"study":"A Retrieval-Augmented Generation Framework for Traditional Chinese Medicine Herb Recommendation Using Symptom-Focused and Ingredient-Based Embeddings","year":2026,"doi":"10.65205/jcct.2026.e3516","url":"https://doi.org/10.65205/jcct.2026.e3516","population":"rag accuracy tasks","intervention_or_exposure":"Retrieval-Augmented Generation Framework","comparator":"not extracted","endpoint":"not extracted","effect":"not extracted","risk_of_bias":"not appraised in public sidecar","directness":"primary","excerpt":"The baseline LLM demonstrated strong performance across multiple metrics, including accuracy (0.1900) and NDCG@5 (0.1475), reflecting substantial pre-trained medical knowledge.","source_id":"source_1","support_kind":"candidate_source_row"},{"study":"Evaluating Retrieval-Augmented Generation Variants for Natural Language-Based SQL and API Call Generation","year":2026,"doi":"10.48550/arxiv.2602.07086","url":"https://doi.org/10.48550/arxiv.2602.07086","population":"combined","intervention_or_exposure":"RAG","comparator":"not extracted","endpoint":"not extracted","effect":"not extracted","risk_of_bias":"not appraised in public sidecar","directness":"primary","excerpt":"Critically, CoRAG proves most robust in hybrid documentation settings, achieving statistically significant improvements in the combined task (10.29% exact match vs. 7.45% for standard RAG), driven primarily by superior SQL generation performance (15.32% vs. 11.56%).","source_id":"source_2","support_kind":"candidate_source_row"},{"study":"A retrieval-augmented generation large language model framework for accurate dementia identification from electronic health records","year":2026,"doi":"10.64898/2026.01.24.26344477","url":"https://doi.org/10.64898/2026.01.24.26344477","population":"rag F1 tasks","intervention_or_exposure":"RAG","comparator":"not extracted","endpoint":"not extracted","effect":"not extracted","risk_of_bias":"not appraised in public sidecar","directness":"primary","excerpt":"ResultsThe RAG-based classifier achieved the highest performance (F1=0.933, sensitivity=91.1%, PPV=95.5%) compared to rule-based (F1=0.823, sensitivity=81.1%, PPV=83.5%) and keyword-filtered LLM (F1=0.903, sensitivity=91.7%, PPV=88.6%).","source_id":"source_3","support_kind":"candidate_source_row"},{"study":"Integrating Dense, Sparse, and Graph-Based Approaches in Financial Data Analysis for a Retrieval-Augmented Generation Framework","year":2026,"doi":"10.1109/acdsa67686.2026.11467963","url":"https://doi.org/10.1109/acdsa67686.2026.11467963","population":"rag recall tasks","intervention_or_exposure":"Integrating Dense, Sparse, and Graph-Based Approaches","comparator":"not extracted","endpoint":"not extracted","effect":"not extracted","risk_of_bias":"not appraised in public sidecar","directness":"primary","excerpt":"Results show that integrating a graph-based retriever improved context recall by 63%, answer correctness by 31%, and overall performance by 12% compared to flattened text retrieval.","source_id":"source_4","support_kind":"candidate_source_row"},{"study":"Improving Retrieval-Augmented Generation Performance Using the MAF-RAG Architecture, EVR–VOR Vector Retrieval, and Multi-Agent Fallback Reasoning","year":2026,"doi":"10.30871/jaic.v10i1.11738","url":"https://doi.org/10.30871/jaic.v10i1.11738","population":"rag F1 tasks","intervention_or_exposure":"RAG","comparator":"not extracted","endpoint":"not extracted","effect":"not extracted","risk_of_bias":"not appraised in public sidecar","directness":"primary","excerpt":"The results show that the proposed MAF-RAG significantly outperforms the baseline system, achieving a mean F1-score of 0.556, an improvement of 18.8% over the Enhanced Baseline (mean F1-score = 0.469) and a 70.0% improvement over the Legacy Baseline (mean F1-score = 0.327).","source_id":"source_5","support_kind":"candidate_source_row"}]},{"claim_id":"claim_3","claim":"3 of 5 selected receipts are direction-bearing for the selected source contexts; 0 receipt(s) are null/mixed and 2 are context/model only. This is a bounded source-literature signal, not a pooled effect.","citation_support":[],"candidate_sources":[{"study":"A Retrieval-Augmented Generation Framework for Traditional Chinese Medicine Herb Recommendation Using Symptom-Focused and Ingredient-Based Embeddings","year":2026,"doi":"10.65205/jcct.2026.e3516","url":"https://doi.org/10.65205/jcct.2026.e3516","population":"rag accuracy tasks","intervention_or_exposure":"Retrieval-Augmented Generation Framework","comparator":"not extracted","endpoint":"not extracted","effect":"not extracted","risk_of_bias":"not appraised in public sidecar","directness":"primary","excerpt":"The baseline LLM demonstrated strong performance across multiple metrics, including accuracy (0.1900) and NDCG@5 (0.1475), reflecting substantial pre-trained medical knowledge.","source_id":"source_1","support_kind":"candidate_source_row"},{"study":"Evaluating Retrieval-Augmented Generation Variants for Natural Language-Based SQL and API Call Generation","year":2026,"doi":"10.48550/arxiv.2602.07086","url":"https://doi.org/10.48550/arxiv.2602.07086","population":"combined","intervention_or_exposure":"RAG","comparator":"not extracted","endpoint":"not extracted","effect":"not extracted","risk_of_bias":"not appraised in public sidecar","directness":"primary","excerpt":"Critically, CoRAG proves most robust in hybrid documentation settings, achieving statistically significant improvements in the combined task (10.29% exact match vs. 7.45% for standard RAG), driven primarily by superior SQL generation performance (15.32% vs. 11.56%).","source_id":"source_2","support_kind":"candidate_source_row"},{"study":"A retrieval-augmented generation large language model framework for accurate dementia identification from electronic health records","year":2026,"doi":"10.64898/2026.01.24.26344477","url":"https://doi.org/10.64898/2026.01.24.26344477","population":"rag F1 tasks","intervention_or_exposure":"RAG","comparator":"not extracted","endpoint":"not extracted","effect":"not extracted","risk_of_bias":"not appraised in public sidecar","directness":"primary","excerpt":"ResultsThe RAG-based classifier achieved the highest performance (F1=0.933, sensitivity=91.1%, PPV=95.5%) compared to rule-based (F1=0.823, sensitivity=81.1%, PPV=83.5%) and keyword-filtered LLM (F1=0.903, sensitivity=91.7%, PPV=88.6%).","source_id":"source_3","support_kind":"candidate_source_row"},{"study":"Integrating Dense, Sparse, and Graph-Based Approaches in Financial Data Analysis for a Retrieval-Augmented Generation Framework","year":2026,"doi":"10.1109/acdsa67686.2026.11467963","url":"https://doi.org/10.1109/acdsa67686.2026.11467963","population":"rag recall tasks","intervention_or_exposure":"Integrating Dense, Sparse, and Graph-Based Approaches","comparator":"not extracted","endpoint":"not extracted","effect":"not extracted","risk_of_bias":"not appraised in public sidecar","directness":"primary","excerpt":"Results show that integrating a graph-based retriever improved context recall by 63%, answer correctness by 31%, and overall performance by 12% compared to flattened text retrieval.","source_id":"source_4","support_kind":"candidate_source_row"},{"study":"Improving Retrieval-Augmented Generation Performance Using the MAF-RAG Architecture, EVR–VOR Vector Retrieval, and Multi-Agent Fallback Reasoning","year":2026,"doi":"10.30871/jaic.v10i1.11738","url":"https://doi.org/10.30871/jaic.v10i1.11738","population":"rag F1 tasks","intervention_or_exposure":"RAG","comparator":"not extracted","endpoint":"not extracted","effect":"not extracted","risk_of_bias":"not appraised in public sidecar","directness":"primary","excerpt":"The results show that the proposed MAF-RAG significantly outperforms the baseline system, achieving a mean F1-score of 0.556, an improvement of 18.8% over the Enhanced Baseline (mean F1-score = 0.469) and a 70.0% improvement over the Legacy Baseline (mean F1-score = 0.327).","source_id":"source_5","support_kind":"candidate_source_row"}]},{"claim_id":"claim_4","claim":"This receipt-backed scoping note has one bounded signal: retrieval augmented generation shows policy/exposure estimates plus separate descriptive evidence across this 5-source primary bundle (2026-2026). Evidence role grouping: direction-bearing receipts: 3; null/mixed metric-scope caveat receipts: 0; context/antecedent/model receipts: 2 excluded from effect support. The source facts cover 4 population/setting context(s) and 3 policy/exposure/practice context(s), so this is a scoping signal about where settings/designs diverge, without establishing a causal, policy-prescriptive, market-generalized, or pooled econometric claim. Population/setting counts are context descriptors only; they are not weighting, pooling, or aggregation evidence. The listed estimates remain source-specific across metrics and settings; they are not pooled or averaged. This is a separated policy/setting map, not a unified pooled economics claim. Named setting scope includes combined, rag F1 tasks, rag accuracy tasks, and rag recall tasks. Within-vs-across outcome rule: direction-bearing rows are only compared within the selected source contexts; unrelated receipt families are not treated as one outcome. Concrete contrast: directional association: Evaluating Retrieval-Augmented Generation Variants for Natural Language-Based SQL and API Call Generation: Critically, CoRAG proves most robust in hybrid documentation settings, achieving statistically significant...; descriptive/modeling: A Retrieval-Augmented Generation Framework for Traditional Chinese Medicine Herb Recommendation Using Symptom-Focused and Ingredient-Based Embeddings: The baseline LLM demonstrated strong performance across multiple metrics, including accuracy (0.1900) and....","citation_support":[],"candidate_sources":[{"study":"A Retrieval-Augmented Generation Framework for Traditional Chinese Medicine Herb Recommendation Using Symptom-Focused and Ingredient-Based Embeddings","year":2026,"doi":"10.65205/jcct.2026.e3516","url":"https://doi.org/10.65205/jcct.2026.e3516","population":"rag accuracy tasks","intervention_or_exposure":"Retrieval-Augmented Generation Framework","comparator":"not extracted","endpoint":"not extracted","effect":"not extracted","risk_of_bias":"not appraised in public sidecar","directness":"primary","excerpt":"The baseline LLM demonstrated strong performance across multiple metrics, including accuracy (0.1900) and NDCG@5 (0.1475), reflecting substantial pre-trained medical knowledge.","source_id":"source_1","support_kind":"candidate_source_row"},{"study":"Evaluating Retrieval-Augmented Generation Variants for Natural Language-Based SQL and API Call Generation","year":2026,"doi":"10.48550/arxiv.2602.07086","url":"https://doi.org/10.48550/arxiv.2602.07086","population":"combined","intervention_or_exposure":"RAG","comparator":"not extracted","endpoint":"not extracted","effect":"not extracted","risk_of_bias":"not appraised in public sidecar","directness":"primary","excerpt":"Critically, CoRAG proves most robust in hybrid documentation settings, achieving statistically significant improvements in the combined task (10.29% exact match vs. 7.45% for standard RAG), driven primarily by superior SQL generation performance (15.32% vs. 11.56%).","source_id":"source_2","support_kind":"candidate_source_row"},{"study":"A retrieval-augmented generation large language model framework for accurate dementia identification from electronic health records","year":2026,"doi":"10.64898/2026.01.24.26344477","url":"https://doi.org/10.64898/2026.01.24.26344477","population":"rag F1 tasks","intervention_or_exposure":"RAG","comparator":"not extracted","endpoint":"not extracted","effect":"not extracted","risk_of_bias":"not appraised in public sidecar","directness":"primary","excerpt":"ResultsThe RAG-based classifier achieved the highest performance (F1=0.933, sensitivity=91.1%, PPV=95.5%) compared to rule-based (F1=0.823, sensitivity=81.1%, PPV=83.5%) and keyword-filtered LLM (F1=0.903, sensitivity=91.7%, PPV=88.6%).","source_id":"source_3","support_kind":"candidate_source_row"},{"study":"Integrating Dense, Sparse, and Graph-Based Approaches in Financial Data Analysis for a Retrieval-Augmented Generation Framework","year":2026,"doi":"10.1109/acdsa67686.2026.11467963","url":"https://doi.org/10.1109/acdsa67686.2026.11467963","population":"rag recall tasks","intervention_or_exposure":"Integrating Dense, Sparse, and Graph-Based Approaches","comparator":"not extracted","endpoint":"not extracted","effect":"not extracted","risk_of_bias":"not appraised in public sidecar","directness":"primary","excerpt":"Results show that integrating a graph-based retriever improved context recall by 63%, answer correctness by 31%, and overall performance by 12% compared to flattened text retrieval.","source_id":"source_4","support_kind":"candidate_source_row"},{"study":"Improving Retrieval-Augmented Generation Performance Using the MAF-RAG Architecture, EVR–VOR Vector Retrieval, and Multi-Agent Fallback Reasoning","year":2026,"doi":"10.30871/jaic.v10i1.11738","url":"https://doi.org/10.30871/jaic.v10i1.11738","population":"rag F1 tasks","intervention_or_exposure":"RAG","comparator":"not extracted","endpoint":"not extracted","effect":"not extracted","risk_of_bias":"not appraised in public sidecar","directness":"primary","excerpt":"The results show that the proposed MAF-RAG significantly outperforms the baseline system, achieving a mean F1-score of 0.556, an improvement of 18.8% over the Enhanced Baseline (mean F1-score = 0.469) and a 70.0% improvement over the Legacy Baseline (mean F1-score = 0.327).","source_id":"source_5","support_kind":"candidate_source_row"}]},{"claim_id":"claim_5","claim":"Role definitions: direction-bearing rows carry metric-specific effect or association text; null/mixed rows carry rejected or non-convergent metric evidence; context/model rows rank, model, or contextualize adjacent constructs. Interpretation: keep these rows separate; do not pool them or treat antecedent/modeling rows as the same estimand.","citation_support":[],"candidate_sources":[{"study":"A Retrieval-Augmented Generation Framework for Traditional Chinese Medicine Herb Recommendation Using Symptom-Focused and Ingredient-Based Embeddings","year":2026,"doi":"10.65205/jcct.2026.e3516","url":"https://doi.org/10.65205/jcct.2026.e3516","population":"rag accuracy tasks","intervention_or_exposure":"Retrieval-Augmented Generation Framework","comparator":"not extracted","endpoint":"not extracted","effect":"not extracted","risk_of_bias":"not appraised in public sidecar","directness":"primary","excerpt":"The baseline LLM demonstrated strong performance across multiple metrics, including accuracy (0.1900) and NDCG@5 (0.1475), reflecting substantial pre-trained medical knowledge.","source_id":"source_1","support_kind":"candidate_source_row"},{"study":"Evaluating Retrieval-Augmented Generation Variants for Natural Language-Based SQL and API Call Generation","year":2026,"doi":"10.48550/arxiv.2602.07086","url":"https://doi.org/10.48550/arxiv.2602.07086","population":"combined","intervention_or_exposure":"RAG","comparator":"not extracted","endpoint":"not extracted","effect":"not extracted","risk_of_bias":"not appraised in public sidecar","directness":"primary","excerpt":"Critically, CoRAG proves most robust in hybrid documentation settings, achieving statistically significant improvements in the combined task (10.29% exact match vs. 7.45% for standard RAG), driven primarily by superior SQL generation performance (15.32% vs. 11.56%).","source_id":"source_2","support_kind":"candidate_source_row"},{"study":"A retrieval-augmented generation large language model framework for accurate dementia identification from electronic health records","year":2026,"doi":"10.64898/2026.01.24.26344477","url":"https://doi.org/10.64898/2026.01.24.26344477","population":"rag F1 tasks","intervention_or_exposure":"RAG","comparator":"not extracted","endpoint":"not extracted","effect":"not extracted","risk_of_bias":"not appraised in public sidecar","directness":"primary","excerpt":"ResultsThe RAG-based classifier achieved the highest performance (F1=0.933, sensitivity=91.1%, PPV=95.5%) compared to rule-based (F1=0.823, sensitivity=81.1%, PPV=83.5%) and keyword-filtered LLM (F1=0.903, sensitivity=91.7%, PPV=88.6%).","source_id":"source_3","support_kind":"candidate_source_row"},{"study":"Integrating Dense, Sparse, and Graph-Based Approaches in Financial Data Analysis for a Retrieval-Augmented Generation Framework","year":2026,"doi":"10.1109/acdsa67686.2026.11467963","url":"https://doi.org/10.1109/acdsa67686.2026.11467963","population":"rag recall tasks","intervention_or_exposure":"Integrating Dense, Sparse, and Graph-Based Approaches","comparator":"not extracted","endpoint":"not extracted","effect":"not extracted","risk_of_bias":"not appraised in public sidecar","directness":"primary","excerpt":"Results show that integrating a graph-based retriever improved context recall by 63%, answer correctness by 31%, and overall performance by 12% compared to flattened text retrieval.","source_id":"source_4","support_kind":"candidate_source_row"},{"study":"Improving Retrieval-Augmented Generation Performance Using the MAF-RAG Architecture, EVR–VOR Vector Retrieval, and Multi-Agent Fallback Reasoning","year":2026,"doi":"10.30871/jaic.v10i1.11738","url":"https://doi.org/10.30871/jaic.v10i1.11738","population":"rag F1 tasks","intervention_or_exposure":"RAG","comparator":"not extracted","endpoint":"not extracted","effect":"not extracted","risk_of_bias":"not appraised in public sidecar","directness":"primary","excerpt":"The results show that the proposed MAF-RAG significantly outperforms the baseline system, achieving a mean F1-score of 0.556, an improvement of 18.8% over the Enhanced Baseline (mean F1-score = 0.469) and a 70.0% improvement over the Legacy Baseline (mean F1-score = 0.327).","source_id":"source_5","support_kind":"candidate_source_row"}]},{"claim_id":"claim_6","claim":"Matrix guard: effect-bearing rows below are metric-specific source facts, not a pooled comparison; context-only rows are excluded from effect support.","citation_support":[],"candidate_sources":[{"study":"A Retrieval-Augmented Generation Framework for Traditional Chinese Medicine Herb Recommendation Using Symptom-Focused and Ingredient-Based Embeddings","year":2026,"doi":"10.65205/jcct.2026.e3516","url":"https://doi.org/10.65205/jcct.2026.e3516","population":"rag accuracy tasks","intervention_or_exposure":"Retrieval-Augmented Generation Framework","comparator":"not extracted","endpoint":"not extracted","effect":"not extracted","risk_of_bias":"not appraised in public sidecar","directness":"primary","excerpt":"The baseline LLM demonstrated strong performance across multiple metrics, including accuracy (0.1900) and NDCG@5 (0.1475), reflecting substantial pre-trained medical knowledge.","source_id":"source_1","support_kind":"candidate_source_row"},{"study":"Evaluating Retrieval-Augmented Generation Variants for Natural Language-Based SQL and API Call Generation","year":2026,"doi":"10.48550/arxiv.2602.07086","url":"https://doi.org/10.48550/arxiv.2602.07086","population":"combined","intervention_or_exposure":"RAG","comparator":"not extracted","endpoint":"not extracted","effect":"not extracted","risk_of_bias":"not appraised in public sidecar","directness":"primary","excerpt":"Critically, CoRAG proves most robust in hybrid documentation settings, achieving statistically significant improvements in the combined task (10.29% exact match vs. 7.45% for standard RAG), driven primarily by superior SQL generation performance (15.32% vs. 11.56%).","source_id":"source_2","support_kind":"candidate_source_row"},{"study":"A retrieval-augmented generation large language model framework for accurate dementia identification from electronic health records","year":2026,"doi":"10.64898/2026.01.24.26344477","url":"https://doi.org/10.64898/2026.01.24.26344477","population":"rag F1 tasks","intervention_or_exposure":"RAG","comparator":"not extracted","endpoint":"not extracted","effect":"not extracted","risk_of_bias":"not appraised in public sidecar","directness":"primary","excerpt":"ResultsThe RAG-based classifier achieved the highest performance (F1=0.933, sensitivity=91.1%, PPV=95.5%) compared to rule-based (F1=0.823, sensitivity=81.1%, PPV=83.5%) and keyword-filtered LLM (F1=0.903, sensitivity=91.7%, PPV=88.6%).","source_id":"source_3","support_kind":"candidate_source_row"},{"study":"Integrating Dense, Sparse, and Graph-Based Approaches in Financial Data Analysis for a Retrieval-Augmented Generation Framework","year":2026,"doi":"10.1109/acdsa67686.2026.11467963","url":"https://doi.org/10.1109/acdsa67686.2026.11467963","population":"rag recall tasks","intervention_or_exposure":"Integrating Dense, Sparse, and Graph-Based Approaches","comparator":"not extracted","endpoint":"not extracted","effect":"not extracted","risk_of_bias":"not appraised in public sidecar","directness":"primary","excerpt":"Results show that integrating a graph-based retriever improved context recall by 63%, answer correctness by 31%, and overall performance by 12% compared to flattened text retrieval.","source_id":"source_4","support_kind":"candidate_source_row"},{"study":"Improving Retrieval-Augmented Generation Performance Using the MAF-RAG Architecture, EVR–VOR Vector Retrieval, and Multi-Agent Fallback Reasoning","year":2026,"doi":"10.30871/jaic.v10i1.11738","url":"https://doi.org/10.30871/jaic.v10i1.11738","population":"rag F1 tasks","intervention_or_exposure":"RAG","comparator":"not extracted","endpoint":"not extracted","effect":"not extracted","risk_of_bias":"not appraised in public sidecar","directness":"primary","excerpt":"The results show that the proposed MAF-RAG significantly outperforms the baseline system, achieving a mean F1-score of 0.556, an improvement of 18.8% over the Enhanced Baseline (mean F1-score = 0.469) and a 70.0% improvement over the Legacy Baseline (mean F1-score = 0.327).","source_id":"source_5","support_kind":"candidate_source_row"}]},{"claim_id":"claim_7","claim":"| Outcome family | Receipt | Evidence role | Population/setting | Metric | Extracted finding |","citation_support":[],"candidate_sources":[{"study":"A Retrieval-Augmented Generation Framework for Traditional Chinese Medicine Herb Recommendation Using Symptom-Focused and Ingredient-Based Embeddings","year":2026,"doi":"10.65205/jcct.2026.e3516","url":"https://doi.org/10.65205/jcct.2026.e3516","population":"rag accuracy tasks","intervention_or_exposure":"Retrieval-Augmented Generation Framework","comparator":"not extracted","endpoint":"not extracted","effect":"not extracted","risk_of_bias":"not appraised in public sidecar","directness":"primary","excerpt":"The baseline LLM demonstrated strong performance across multiple metrics, including accuracy (0.1900) and NDCG@5 (0.1475), reflecting substantial pre-trained medical knowledge.","source_id":"source_1","support_kind":"candidate_source_row"},{"study":"Evaluating Retrieval-Augmented Generation Variants for Natural Language-Based SQL and API Call Generation","year":2026,"doi":"10.48550/arxiv.2602.07086","url":"https://doi.org/10.48550/arxiv.2602.07086","population":"combined","intervention_or_exposure":"RAG","comparator":"not extracted","endpoint":"not extracted","effect":"not extracted","risk_of_bias":"not appraised in public sidecar","directness":"primary","excerpt":"Critically, CoRAG proves most robust in hybrid documentation settings, achieving statistically significant improvements in the combined task (10.29% exact match vs. 7.45% for standard RAG), driven primarily by superior SQL generation performance (15.32% vs. 11.56%).","source_id":"source_2","support_kind":"candidate_source_row"},{"study":"A retrieval-augmented generation large language model framework for accurate dementia identification from electronic health records","year":2026,"doi":"10.64898/2026.01.24.26344477","url":"https://doi.org/10.64898/2026.01.24.26344477","population":"rag F1 tasks","intervention_or_exposure":"RAG","comparator":"not extracted","endpoint":"not extracted","effect":"not extracted","risk_of_bias":"not appraised in public sidecar","directness":"primary","excerpt":"ResultsThe RAG-based classifier achieved the highest performance (F1=0.933, sensitivity=91.1%, PPV=95.5%) compared to rule-based (F1=0.823, sensitivity=81.1%, PPV=83.5%) and keyword-filtered LLM (F1=0.903, sensitivity=91.7%, PPV=88.6%).","source_id":"source_3","support_kind":"candidate_source_row"},{"study":"Integrating Dense, Sparse, and Graph-Based Approaches in Financial Data Analysis for a Retrieval-Augmented Generation Framework","year":2026,"doi":"10.1109/acdsa67686.2026.11467963","url":"https://doi.org/10.1109/acdsa67686.2026.11467963","population":"rag recall tasks","intervention_or_exposure":"Integrating Dense, Sparse, and Graph-Based Approaches","comparator":"not extracted","endpoint":"not extracted","effect":"not extracted","risk_of_bias":"not appraised in public sidecar","directness":"primary","excerpt":"Results show that integrating a graph-based retriever improved context recall by 63%, answer correctness by 31%, and overall performance by 12% compared to flattened text retrieval.","source_id":"source_4","support_kind":"candidate_source_row"},{"study":"Improving Retrieval-Augmented Generation Performance Using the MAF-RAG Architecture, EVR–VOR Vector Retrieval, and Multi-Agent Fallback Reasoning","year":2026,"doi":"10.30871/jaic.v10i1.11738","url":"https://doi.org/10.30871/jaic.v10i1.11738","population":"rag F1 tasks","intervention_or_exposure":"RAG","comparator":"not extracted","endpoint":"not extracted","effect":"not extracted","risk_of_bias":"not appraised in public sidecar","directness":"primary","excerpt":"The results show that the proposed MAF-RAG significantly outperforms the baseline system, achieving a mean F1-score of 0.556, an improvement of 18.8% over the Enhanced Baseline (mean F1-score = 0.469) and a 70.0% improvement over the Legacy Baseline (mean F1-score = 0.327).","source_id":"source_5","support_kind":"candidate_source_row"}]},{"claim_id":"claim_8","claim":"| Outcome family | Receipt | Evidence role | Population/setting | Metric | Extracted finding |","citation_support":[],"candidate_sources":[{"study":"A Retrieval-Augmented Generation Framework for Traditional Chinese Medicine Herb Recommendation Using Symptom-Focused and Ingredient-Based Embeddings","year":2026,"doi":"10.65205/jcct.2026.e3516","url":"https://doi.org/10.65205/jcct.2026.e3516","population":"rag accuracy tasks","intervention_or_exposure":"Retrieval-Augmented Generation Framework","comparator":"not extracted","endpoint":"not extracted","effect":"not extracted","risk_of_bias":"not appraised in public sidecar","directness":"primary","excerpt":"The baseline LLM demonstrated strong performance across multiple metrics, including accuracy (0.1900) and NDCG@5 (0.1475), reflecting substantial pre-trained medical knowledge.","source_id":"source_1","support_kind":"candidate_source_row"},{"study":"Evaluating Retrieval-Augmented Generation Variants for Natural Language-Based SQL and API Call Generation","year":2026,"doi":"10.48550/arxiv.2602.07086","url":"https://doi.org/10.48550/arxiv.2602.07086","population":"combined","intervention_or_exposure":"RAG","comparator":"not extracted","endpoint":"not extracted","effect":"not extracted","risk_of_bias":"not appraised in public sidecar","directness":"primary","excerpt":"Critically, CoRAG proves most robust in hybrid documentation settings, achieving statistically significant improvements in the combined task (10.29% exact match vs. 7.45% for standard RAG), driven primarily by superior SQL generation performance (15.32% vs. 11.56%).","source_id":"source_2","support_kind":"candidate_source_row"},{"study":"A retrieval-augmented generation large language model framework for accurate dementia identification from electronic health records","year":2026,"doi":"10.64898/2026.01.24.26344477","url":"https://doi.org/10.64898/2026.01.24.26344477","population":"rag F1 tasks","intervention_or_exposure":"RAG","comparator":"not extracted","endpoint":"not extracted","effect":"not extracted","risk_of_bias":"not appraised in public sidecar","directness":"primary","excerpt":"ResultsThe RAG-based classifier achieved the highest performance (F1=0.933, sensitivity=91.1%, PPV=95.5%) compared to rule-based (F1=0.823, sensitivity=81.1%, PPV=83.5%) and keyword-filtered LLM (F1=0.903, sensitivity=91.7%, PPV=88.6%).","source_id":"source_3","support_kind":"candidate_source_row"},{"study":"Integrating Dense, Sparse, and Graph-Based Approaches in Financial Data Analysis for a Retrieval-Augmented Generation Framework","year":2026,"doi":"10.1109/acdsa67686.2026.11467963","url":"https://doi.org/10.1109/acdsa67686.2026.11467963","population":"rag recall tasks","intervention_or_exposure":"Integrating Dense, Sparse, and Graph-Based Approaches","comparator":"not extracted","endpoint":"not extracted","effect":"not extracted","risk_of_bias":"not appraised in public sidecar","directness":"primary","excerpt":"Results show that integrating a graph-based retriever improved context recall by 63%, answer correctness by 31%, and overall performance by 12% compared to flattened text retrieval.","source_id":"source_4","support_kind":"candidate_source_row"},{"study":"Improving Retrieval-Augmented Generation Performance Using the MAF-RAG Architecture, EVR–VOR Vector Retrieval, and Multi-Agent Fallback Reasoning","year":2026,"doi":"10.30871/jaic.v10i1.11738","url":"https://doi.org/10.30871/jaic.v10i1.11738","population":"rag F1 tasks","intervention_or_exposure":"RAG","comparator":"not extracted","endpoint":"not extracted","effect":"not extracted","risk_of_bias":"not appraised in public sidecar","directness":"primary","excerpt":"The results show that the proposed MAF-RAG significantly outperforms the baseline system, achieving a mean F1-score of 0.556, an improvement of 18.8% over the Enhanced Baseline (mean F1-score = 0.469) and a 70.0% improvement over the Legacy Baseline (mean F1-score = 0.327).","source_id":"source_5","support_kind":"candidate_source_row"}]},{"claim_id":"claim_9","claim":"Audit note: effect-bearing rows stay metric-specific; context-only rows are excluded from effect support; role counts below keep direction-bearing, null/mixed metric-scope caveat, and context-only receipts separate.","citation_support":[],"candidate_sources":[{"study":"A Retrieval-Augmented Generation Framework for Traditional Chinese Medicine Herb Recommendation Using Symptom-Focused and Ingredient-Based Embeddings","year":2026,"doi":"10.65205/jcct.2026.e3516","url":"https://doi.org/10.65205/jcct.2026.e3516","population":"rag accuracy tasks","intervention_or_exposure":"Retrieval-Augmented Generation Framework","comparator":"not extracted","endpoint":"not extracted","effect":"not extracted","risk_of_bias":"not appraised in public sidecar","directness":"primary","excerpt":"The baseline LLM demonstrated strong performance across multiple metrics, including accuracy (0.1900) and NDCG@5 (0.1475), reflecting substantial pre-trained medical knowledge.","source_id":"source_1","support_kind":"candidate_source_row"},{"study":"Evaluating Retrieval-Augmented Generation Variants for Natural Language-Based SQL and API Call Generation","year":2026,"doi":"10.48550/arxiv.2602.07086","url":"https://doi.org/10.48550/arxiv.2602.07086","population":"combined","intervention_or_exposure":"RAG","comparator":"not extracted","endpoint":"not extracted","effect":"not extracted","risk_of_bias":"not appraised in public sidecar","directness":"primary","excerpt":"Critically, CoRAG proves most robust in hybrid documentation settings, achieving statistically significant improvements in the combined task (10.29% exact match vs. 7.45% for standard RAG), driven primarily by superior SQL generation performance (15.32% vs. 11.56%).","source_id":"source_2","support_kind":"candidate_source_row"},{"study":"A retrieval-augmented generation large language model framework for accurate dementia identification from electronic health records","year":2026,"doi":"10.64898/2026.01.24.26344477","url":"https://doi.org/10.64898/2026.01.24.26344477","population":"rag F1 tasks","intervention_or_exposure":"RAG","comparator":"not extracted","endpoint":"not extracted","effect":"not extracted","risk_of_bias":"not appraised in public sidecar","directness":"primary","excerpt":"ResultsThe RAG-based classifier achieved the highest performance (F1=0.933, sensitivity=91.1%, PPV=95.5%) compared to rule-based (F1=0.823, sensitivity=81.1%, PPV=83.5%) and keyword-filtered LLM (F1=0.903, sensitivity=91.7%, PPV=88.6%).","source_id":"source_3","support_kind":"candidate_source_row"},{"study":"Integrating Dense, Sparse, and Graph-Based Approaches in Financial Data Analysis for a Retrieval-Augmented Generation Framework","year":2026,"doi":"10.1109/acdsa67686.2026.11467963","url":"https://doi.org/10.1109/acdsa67686.2026.11467963","population":"rag recall tasks","intervention_or_exposure":"Integrating Dense, Sparse, and Graph-Based Approaches","comparator":"not extracted","endpoint":"not extracted","effect":"not extracted","risk_of_bias":"not appraised in public sidecar","directness":"primary","excerpt":"Results show that integrating a graph-based retriever improved context recall by 63%, answer correctness by 31%, and overall performance by 12% compared to flattened text retrieval.","source_id":"source_4","support_kind":"candidate_source_row"},{"study":"Improving Retrieval-Augmented Generation Performance Using the MAF-RAG Architecture, EVR–VOR Vector Retrieval, and Multi-Agent Fallback Reasoning","year":2026,"doi":"10.30871/jaic.v10i1.11738","url":"https://doi.org/10.30871/jaic.v10i1.11738","population":"rag F1 tasks","intervention_or_exposure":"RAG","comparator":"not extracted","endpoint":"not extracted","effect":"not extracted","risk_of_bias":"not appraised in public sidecar","directness":"primary","excerpt":"The results show that the proposed MAF-RAG significantly outperforms the baseline system, achieving a mean F1-score of 0.556, an improvement of 18.8% over the Enhanced Baseline (mean F1-score = 0.469) and a 70.0% improvement over the Legacy Baseline (mean F1-score = 0.327).","source_id":"source_5","support_kind":"candidate_source_row"}]},{"claim_id":"claim_10","claim":"Evidence role summary: direction-bearing receipts: 3; null/mixed metric-scope caveat receipts: 0; context/antecedent/model receipts: 2 excluded from effect support.","citation_support":[],"candidate_sources":[{"study":"A Retrieval-Augmented Generation Framework for Traditional Chinese Medicine Herb Recommendation Using Symptom-Focused and Ingredient-Based Embeddings","year":2026,"doi":"10.65205/jcct.2026.e3516","url":"https://doi.org/10.65205/jcct.2026.e3516","population":"rag accuracy tasks","intervention_or_exposure":"Retrieval-Augmented Generation Framework","comparator":"not extracted","endpoint":"not extracted","effect":"not extracted","risk_of_bias":"not appraised in public sidecar","directness":"primary","excerpt":"The baseline LLM demonstrated strong performance across multiple metrics, including accuracy (0.1900) and NDCG@5 (0.1475), reflecting substantial pre-trained medical knowledge.","source_id":"source_1","support_kind":"candidate_source_row"},{"study":"Evaluating Retrieval-Augmented Generation Variants for Natural Language-Based SQL and API Call Generation","year":2026,"doi":"10.48550/arxiv.2602.07086","url":"https://doi.org/10.48550/arxiv.2602.07086","population":"combined","intervention_or_exposure":"RAG","comparator":"not extracted","endpoint":"not extracted","effect":"not extracted","risk_of_bias":"not appraised in public sidecar","directness":"primary","excerpt":"Critically, CoRAG proves most robust in hybrid documentation settings, achieving statistically significant improvements in the combined task (10.29% exact match vs. 7.45% for standard RAG), driven primarily by superior SQL generation performance (15.32% vs. 11.56%).","source_id":"source_2","support_kind":"candidate_source_row"},{"study":"A retrieval-augmented generation large language model framework for accurate dementia identification from electronic health records","year":2026,"doi":"10.64898/2026.01.24.26344477","url":"https://doi.org/10.64898/2026.01.24.26344477","population":"rag F1 tasks","intervention_or_exposure":"RAG","comparator":"not extracted","endpoint":"not extracted","effect":"not extracted","risk_of_bias":"not appraised in public sidecar","directness":"primary","excerpt":"ResultsThe RAG-based classifier achieved the highest performance (F1=0.933, sensitivity=91.1%, PPV=95.5%) compared to rule-based (F1=0.823, sensitivity=81.1%, PPV=83.5%) and keyword-filtered LLM (F1=0.903, sensitivity=91.7%, PPV=88.6%).","source_id":"source_3","support_kind":"candidate_source_row"},{"study":"Integrating Dense, Sparse, and Graph-Based Approaches in Financial Data Analysis for a Retrieval-Augmented Generation Framework","year":2026,"doi":"10.1109/acdsa67686.2026.11467963","url":"https://doi.org/10.1109/acdsa67686.2026.11467963","population":"rag recall tasks","intervention_or_exposure":"Integrating Dense, Sparse, and Graph-Based Approaches","comparator":"not extracted","endpoint":"not extracted","effect":"not extracted","risk_of_bias":"not appraised in public sidecar","directness":"primary","excerpt":"Results show that integrating a graph-based retriever improved context recall by 63%, answer correctness by 31%, and overall performance by 12% compared to flattened text retrieval.","source_id":"source_4","support_kind":"candidate_source_row"},{"study":"Improving Retrieval-Augmented Generation Performance Using the MAF-RAG Architecture, EVR–VOR Vector Retrieval, and Multi-Agent Fallback Reasoning","year":2026,"doi":"10.30871/jaic.v10i1.11738","url":"https://doi.org/10.30871/jaic.v10i1.11738","population":"rag F1 tasks","intervention_or_exposure":"RAG","comparator":"not extracted","endpoint":"not extracted","effect":"not extracted","risk_of_bias":"not appraised in public sidecar","directness":"primary","excerpt":"The results show that the proposed MAF-RAG significantly outperforms the baseline system, achieving a mean F1-score of 0.556, an improvement of 18.8% over the Enhanced Baseline (mean F1-score = 0.469) and a 70.0% improvement over the Legacy Baseline (mean F1-score = 0.327).","source_id":"source_5","support_kind":"candidate_source_row"}]},{"claim_id":"claim_11","claim":"Specific moderators in this bundle are population/indication (combined; rag F1 tasks; rag accuracy tasks; rag recall tasks), study design/evidence type (primary).","citation_support":[],"candidate_sources":[{"study":"A Retrieval-Augmented Generation Framework for Traditional Chinese Medicine Herb Recommendation Using Symptom-Focused and Ingredient-Based Embeddings","year":2026,"doi":"10.65205/jcct.2026.e3516","url":"https://doi.org/10.65205/jcct.2026.e3516","population":"rag accuracy tasks","intervention_or_exposure":"Retrieval-Augmented Generation Framework","comparator":"not extracted","endpoint":"not extracted","effect":"not extracted","risk_of_bias":"not appraised in public sidecar","directness":"primary","excerpt":"The baseline LLM demonstrated strong performance across multiple metrics, including accuracy (0.1900) and NDCG@5 (0.1475), reflecting substantial pre-trained medical knowledge.","source_id":"source_1","support_kind":"candidate_source_row"},{"study":"Evaluating Retrieval-Augmented Generation Variants for Natural Language-Based SQL and API Call Generation","year":2026,"doi":"10.48550/arxiv.2602.07086","url":"https://doi.org/10.48550/arxiv.2602.07086","population":"combined","intervention_or_exposure":"RAG","comparator":"not extracted","endpoint":"not extracted","effect":"not extracted","risk_of_bias":"not appraised in public sidecar","directness":"primary","excerpt":"Critically, CoRAG proves most robust in hybrid documentation settings, achieving statistically significant improvements in the combined task (10.29% exact match vs. 7.45% for standard RAG), driven primarily by superior SQL generation performance (15.32% vs. 11.56%).","source_id":"source_2","support_kind":"candidate_source_row"},{"study":"A retrieval-augmented generation large language model framework for accurate dementia identification from electronic health records","year":2026,"doi":"10.64898/2026.01.24.26344477","url":"https://doi.org/10.64898/2026.01.24.26344477","population":"rag F1 tasks","intervention_or_exposure":"RAG","comparator":"not extracted","endpoint":"not extracted","effect":"not extracted","risk_of_bias":"not appraised in public sidecar","directness":"primary","excerpt":"ResultsThe RAG-based classifier achieved the highest performance (F1=0.933, sensitivity=91.1%, PPV=95.5%) compared to rule-based (F1=0.823, sensitivity=81.1%, PPV=83.5%) and keyword-filtered LLM (F1=0.903, sensitivity=91.7%, PPV=88.6%).","source_id":"source_3","support_kind":"candidate_source_row"},{"study":"Integrating Dense, Sparse, and Graph-Based Approaches in Financial Data Analysis for a Retrieval-Augmented Generation Framework","year":2026,"doi":"10.1109/acdsa67686.2026.11467963","url":"https://doi.org/10.1109/acdsa67686.2026.11467963","population":"rag recall tasks","intervention_or_exposure":"Integrating Dense, Sparse, and Graph-Based Approaches","comparator":"not extracted","endpoint":"not extracted","effect":"not extracted","risk_of_bias":"not appraised in public sidecar","directness":"primary","excerpt":"Results show that integrating a graph-based retriever improved context recall by 63%, answer correctness by 31%, and overall performance by 12% compared to flattened text retrieval.","source_id":"source_4","support_kind":"candidate_source_row"},{"study":"Improving Retrieval-Augmented Generation Performance Using the MAF-RAG Architecture, EVR–VOR Vector Retrieval, and Multi-Agent Fallback Reasoning","year":2026,"doi":"10.30871/jaic.v10i1.11738","url":"https://doi.org/10.30871/jaic.v10i1.11738","population":"rag F1 tasks","intervention_or_exposure":"RAG","comparator":"not extracted","endpoint":"not extracted","effect":"not extracted","risk_of_bias":"not appraised in public sidecar","directness":"primary","excerpt":"The results show that the proposed MAF-RAG significantly outperforms the baseline system, achieving a mean F1-score of 0.556, an improvement of 18.8% over the Enhanced Baseline (mean F1-score = 0.469) and a 70.0% improvement over the Legacy Baseline (mean F1-score = 0.327).","source_id":"source_5","support_kind":"candidate_source_row"}]},{"claim_id":"claim_12","claim":"Population/settings are separated as receipt context: combined, rag F1 tasks, rag accuracy tasks, and rag recall tasks. The selected receipts group because each carries a fact-level extraction for retrieval augmented generation; they separate by context (other source context) and metric, so they are not interchangeable evidence for one pooled claim.","citation_support":[],"candidate_sources":[{"study":"A Retrieval-Augmented Generation Framework for Traditional Chinese Medicine Herb Recommendation Using Symptom-Focused and Ingredient-Based Embeddings","year":2026,"doi":"10.65205/jcct.2026.e3516","url":"https://doi.org/10.65205/jcct.2026.e3516","population":"rag accuracy tasks","intervention_or_exposure":"Retrieval-Augmented Generation Framework","comparator":"not extracted","endpoint":"not extracted","effect":"not extracted","risk_of_bias":"not appraised in public sidecar","directness":"primary","excerpt":"The baseline LLM demonstrated strong performance across multiple metrics, including accuracy (0.1900) and NDCG@5 (0.1475), reflecting substantial pre-trained medical knowledge.","source_id":"source_1","support_kind":"candidate_source_row"},{"study":"Evaluating Retrieval-Augmented Generation Variants for Natural Language-Based SQL and API Call Generation","year":2026,"doi":"10.48550/arxiv.2602.07086","url":"https://doi.org/10.48550/arxiv.2602.07086","population":"combined","intervention_or_exposure":"RAG","comparator":"not extracted","endpoint":"not extracted","effect":"not extracted","risk_of_bias":"not appraised in public sidecar","directness":"primary","excerpt":"Critically, CoRAG proves most robust in hybrid documentation settings, achieving statistically significant improvements in the combined task (10.29% exact match vs. 7.45% for standard RAG), driven primarily by superior SQL generation performance (15.32% vs. 11.56%).","source_id":"source_2","support_kind":"candidate_source_row"},{"study":"A retrieval-augmented generation large language model framework for accurate dementia identification from electronic health records","year":2026,"doi":"10.64898/2026.01.24.26344477","url":"https://doi.org/10.64898/2026.01.24.26344477","population":"rag F1 tasks","intervention_or_exposure":"RAG","comparator":"not extracted","endpoint":"not extracted","effect":"not extracted","risk_of_bias":"not appraised in public sidecar","directness":"primary","excerpt":"ResultsThe RAG-based classifier achieved the highest performance (F1=0.933, sensitivity=91.1%, PPV=95.5%) compared to rule-based (F1=0.823, sensitivity=81.1%, PPV=83.5%) and keyword-filtered LLM (F1=0.903, sensitivity=91.7%, PPV=88.6%).","source_id":"source_3","support_kind":"candidate_source_row"},{"study":"Integrating Dense, Sparse, and Graph-Based Approaches in Financial Data Analysis for a Retrieval-Augmented Generation Framework","year":2026,"doi":"10.1109/acdsa67686.2026.11467963","url":"https://doi.org/10.1109/acdsa67686.2026.11467963","population":"rag recall tasks","intervention_or_exposure":"Integrating Dense, Sparse, and Graph-Based Approaches","comparator":"not extracted","endpoint":"not extracted","effect":"not extracted","risk_of_bias":"not appraised in public sidecar","directness":"primary","excerpt":"Results show that integrating a graph-based retriever improved context recall by 63%, answer correctness by 31%, and overall performance by 12% compared to flattened text retrieval.","source_id":"source_4","support_kind":"candidate_source_row"},{"study":"Improving Retrieval-Augmented Generation Performance Using the MAF-RAG Architecture, EVR–VOR Vector Retrieval, and Multi-Agent Fallback Reasoning","year":2026,"doi":"10.30871/jaic.v10i1.11738","url":"https://doi.org/10.30871/jaic.v10i1.11738","population":"rag F1 tasks","intervention_or_exposure":"RAG","comparator":"not extracted","endpoint":"not extracted","effect":"not extracted","risk_of_bias":"not appraised in public sidecar","directness":"primary","excerpt":"The results show that the proposed MAF-RAG significantly outperforms the baseline system, achieving a mean F1-score of 0.556, an improvement of 18.8% over the Enhanced Baseline (mean F1-score = 0.469) and a 70.0% improvement over the Legacy Baseline (mean F1-score = 0.327).","source_id":"source_5","support_kind":"candidate_source_row"}]},{"claim_id":"claim_13","claim":"The signal is purely descriptive of source-level direction and scope; it cannot support a causal, policy-prescriptive, or pooled elasticity inference, and pooling across these designs would be inappropriate.","citation_support":[],"candidate_sources":[{"study":"A Retrieval-Augmented Generation Framework for Traditional Chinese Medicine Herb Recommendation Using Symptom-Focused and Ingredient-Based Embeddings","year":2026,"doi":"10.65205/jcct.2026.e3516","url":"https://doi.org/10.65205/jcct.2026.e3516","population":"rag accuracy tasks","intervention_or_exposure":"Retrieval-Augmented Generation Framework","comparator":"not extracted","endpoint":"not extracted","effect":"not extracted","risk_of_bias":"not appraised in public sidecar","directness":"primary","excerpt":"The baseline LLM demonstrated strong performance across multiple metrics, including accuracy (0.1900) and NDCG@5 (0.1475), reflecting substantial pre-trained medical knowledge.","source_id":"source_1","support_kind":"candidate_source_row"},{"study":"Evaluating Retrieval-Augmented Generation Variants for Natural Language-Based SQL and API Call Generation","year":2026,"doi":"10.48550/arxiv.2602.07086","url":"https://doi.org/10.48550/arxiv.2602.07086","population":"combined","intervention_or_exposure":"RAG","comparator":"not extracted","endpoint":"not extracted","effect":"not extracted","risk_of_bias":"not appraised in public sidecar","directness":"primary","excerpt":"Critically, CoRAG proves most robust in hybrid documentation settings, achieving statistically significant improvements in the combined task (10.29% exact match vs. 7.45% for standard RAG), driven primarily by superior SQL generation performance (15.32% vs. 11.56%).","source_id":"source_2","support_kind":"candidate_source_row"},{"study":"A retrieval-augmented generation large language model framework for accurate dementia identification from electronic health records","year":2026,"doi":"10.64898/2026.01.24.26344477","url":"https://doi.org/10.64898/2026.01.24.26344477","population":"rag F1 tasks","intervention_or_exposure":"RAG","comparator":"not extracted","endpoint":"not extracted","effect":"not extracted","risk_of_bias":"not appraised in public sidecar","directness":"primary","excerpt":"ResultsThe RAG-based classifier achieved the highest performance (F1=0.933, sensitivity=91.1%, PPV=95.5%) compared to rule-based (F1=0.823, sensitivity=81.1%, PPV=83.5%) and keyword-filtered LLM (F1=0.903, sensitivity=91.7%, PPV=88.6%).","source_id":"source_3","support_kind":"candidate_source_row"},{"study":"Integrating Dense, Sparse, and Graph-Based Approaches in Financial Data Analysis for a Retrieval-Augmented Generation Framework","year":2026,"doi":"10.1109/acdsa67686.2026.11467963","url":"https://doi.org/10.1109/acdsa67686.2026.11467963","population":"rag recall tasks","intervention_or_exposure":"Integrating Dense, Sparse, and Graph-Based Approaches","comparator":"not extracted","endpoint":"not extracted","effect":"not extracted","risk_of_bias":"not appraised in public sidecar","directness":"primary","excerpt":"Results show that integrating a graph-based retriever improved context recall by 63%, answer correctness by 31%, and overall performance by 12% compared to flattened text retrieval.","source_id":"source_4","support_kind":"candidate_source_row"},{"study":"Improving Retrieval-Augmented Generation Performance Using the MAF-RAG Architecture, EVR–VOR Vector Retrieval, and Multi-Agent Fallback Reasoning","year":2026,"doi":"10.30871/jaic.v10i1.11738","url":"https://doi.org/10.30871/jaic.v10i1.11738","population":"rag F1 tasks","intervention_or_exposure":"RAG","comparator":"not extracted","endpoint":"not extracted","effect":"not extracted","risk_of_bias":"not appraised in public sidecar","directness":"primary","excerpt":"The results show that the proposed MAF-RAG significantly outperforms the baseline system, achieving a mean F1-score of 0.556, an improvement of 18.8% over the Enhanced Baseline (mean F1-score = 0.469) and a 70.0% improvement over the Legacy Baseline (mean F1-score = 0.327).","source_id":"source_5","support_kind":"candidate_source_row"}]},{"claim_id":"claim_14","claim":"Effect-support accounting: 2 of 5 receipt(s) is context/modeling-only and contributes no effect estimate; 3 receipt(s) are direction-bearing and 0 receipt(s) are null/mixed metric-scope caveats.","citation_support":[],"candidate_sources":[{"study":"A Retrieval-Augmented Generation Framework for Traditional Chinese Medicine Herb Recommendation Using Symptom-Focused and Ingredient-Based Embeddings","year":2026,"doi":"10.65205/jcct.2026.e3516","url":"https://doi.org/10.65205/jcct.2026.e3516","population":"rag accuracy tasks","intervention_or_exposure":"Retrieval-Augmented Generation Framework","comparator":"not extracted","endpoint":"not extracted","effect":"not extracted","risk_of_bias":"not appraised in public sidecar","directness":"primary","excerpt":"The baseline LLM demonstrated strong performance across multiple metrics, including accuracy (0.1900) and NDCG@5 (0.1475), reflecting substantial pre-trained medical knowledge.","source_id":"source_1","support_kind":"candidate_source_row"},{"study":"Evaluating Retrieval-Augmented Generation Variants for Natural Language-Based SQL and API Call Generation","year":2026,"doi":"10.48550/arxiv.2602.07086","url":"https://doi.org/10.48550/arxiv.2602.07086","population":"combined","intervention_or_exposure":"RAG","comparator":"not extracted","endpoint":"not extracted","effect":"not extracted","risk_of_bias":"not appraised in public sidecar","directness":"primary","excerpt":"Critically, CoRAG proves most robust in hybrid documentation settings, achieving statistically significant improvements in the combined task (10.29% exact match vs. 7.45% for standard RAG), driven primarily by superior SQL generation performance (15.32% vs. 11.56%).","source_id":"source_2","support_kind":"candidate_source_row"},{"study":"A retrieval-augmented generation large language model framework for accurate dementia identification from electronic health records","year":2026,"doi":"10.64898/2026.01.24.26344477","url":"https://doi.org/10.64898/2026.01.24.26344477","population":"rag F1 tasks","intervention_or_exposure":"RAG","comparator":"not extracted","endpoint":"not extracted","effect":"not extracted","risk_of_bias":"not appraised in public sidecar","directness":"primary","excerpt":"ResultsThe RAG-based classifier achieved the highest performance (F1=0.933, sensitivity=91.1%, PPV=95.5%) compared to rule-based (F1=0.823, sensitivity=81.1%, PPV=83.5%) and keyword-filtered LLM (F1=0.903, sensitivity=91.7%, PPV=88.6%).","source_id":"source_3","support_kind":"candidate_source_row"},{"study":"Integrating Dense, Sparse, and Graph-Based Approaches in Financial Data Analysis for a Retrieval-Augmented Generation Framework","year":2026,"doi":"10.1109/acdsa67686.2026.11467963","url":"https://doi.org/10.1109/acdsa67686.2026.11467963","population":"rag recall tasks","intervention_or_exposure":"Integrating Dense, Sparse, and Graph-Based Approaches","comparator":"not extracted","endpoint":"not extracted","effect":"not extracted","risk_of_bias":"not appraised in public sidecar","directness":"primary","excerpt":"Results show that integrating a graph-based retriever improved context recall by 63%, answer correctness by 31%, and overall performance by 12% compared to flattened text retrieval.","source_id":"source_4","support_kind":"candidate_source_row"},{"study":"Improving Retrieval-Augmented Generation Performance Using the MAF-RAG Architecture, EVR–VOR Vector Retrieval, and Multi-Agent Fallback Reasoning","year":2026,"doi":"10.30871/jaic.v10i1.11738","url":"https://doi.org/10.30871/jaic.v10i1.11738","population":"rag F1 tasks","intervention_or_exposure":"RAG","comparator":"not extracted","endpoint":"not extracted","effect":"not extracted","risk_of_bias":"not appraised in public sidecar","directness":"primary","excerpt":"The results show that the proposed MAF-RAG significantly outperforms the baseline system, achieving a mean F1-score of 0.556, an improvement of 18.8% over the Enhanced Baseline (mean F1-score = 0.469) and a 70.0% improvement over the Legacy Baseline (mean F1-score = 0.327).","source_id":"source_5","support_kind":"candidate_source_row"}]},{"claim_id":"claim_15","claim":"This scoping signal would weaken if the null/mixed metric replicates in matched designs, if direction-bearing rows fail to reproduce within their named metric family, or if context/model rows become the only topic-overlapping receipts.","citation_support":[],"candidate_sources":[{"study":"A Retrieval-Augmented Generation Framework for Traditional Chinese Medicine Herb Recommendation Using Symptom-Focused and Ingredient-Based Embeddings","year":2026,"doi":"10.65205/jcct.2026.e3516","url":"https://doi.org/10.65205/jcct.2026.e3516","population":"rag accuracy tasks","intervention_or_exposure":"Retrieval-Augmented Generation Framework","comparator":"not extracted","endpoint":"not extracted","effect":"not extracted","risk_of_bias":"not appraised in public sidecar","directness":"primary","excerpt":"The baseline LLM demonstrated strong performance across multiple metrics, including accuracy (0.1900) and NDCG@5 (0.1475), reflecting substantial pre-trained medical knowledge.","source_id":"source_1","support_kind":"candidate_source_row"},{"study":"Evaluating Retrieval-Augmented Generation Variants for Natural Language-Based SQL and API Call Generation","year":2026,"doi":"10.48550/arxiv.2602.07086","url":"https://doi.org/10.48550/arxiv.2602.07086","population":"combined","intervention_or_exposure":"RAG","comparator":"not extracted","endpoint":"not extracted","effect":"not extracted","risk_of_bias":"not appraised in public sidecar","directness":"primary","excerpt":"Critically, CoRAG proves most robust in hybrid documentation settings, achieving statistically significant improvements in the combined task (10.29% exact match vs. 7.45% for standard RAG), driven primarily by superior SQL generation performance (15.32% vs. 11.56%).","source_id":"source_2","support_kind":"candidate_source_row"},{"study":"A retrieval-augmented generation large language model framework for accurate dementia identification from electronic health records","year":2026,"doi":"10.64898/2026.01.24.26344477","url":"https://doi.org/10.64898/2026.01.24.26344477","population":"rag F1 tasks","intervention_or_exposure":"RAG","comparator":"not extracted","endpoint":"not extracted","effect":"not extracted","risk_of_bias":"not appraised in public sidecar","directness":"primary","excerpt":"ResultsThe RAG-based classifier achieved the highest performance (F1=0.933, sensitivity=91.1%, PPV=95.5%) compared to rule-based (F1=0.823, sensitivity=81.1%, PPV=83.5%) and keyword-filtered LLM (F1=0.903, sensitivity=91.7%, PPV=88.6%).","source_id":"source_3","support_kind":"candidate_source_row"},{"study":"Integrating Dense, Sparse, and Graph-Based Approaches in Financial Data Analysis for a Retrieval-Augmented Generation Framework","year":2026,"doi":"10.1109/acdsa67686.2026.11467963","url":"https://doi.org/10.1109/acdsa67686.2026.11467963","population":"rag recall tasks","intervention_or_exposure":"Integrating Dense, Sparse, and Graph-Based Approaches","comparator":"not extracted","endpoint":"not extracted","effect":"not extracted","risk_of_bias":"not appraised in public sidecar","directness":"primary","excerpt":"Results show that integrating a graph-based retriever improved context recall by 63%, answer correctness by 31%, and overall performance by 12% compared to flattened text retrieval.","source_id":"source_4","support_kind":"candidate_source_row"},{"study":"Improving Retrieval-Augmented Generation Performance Using the MAF-RAG Architecture, EVR–VOR Vector Retrieval, and Multi-Agent Fallback Reasoning","year":2026,"doi":"10.30871/jaic.v10i1.11738","url":"https://doi.org/10.30871/jaic.v10i1.11738","population":"rag F1 tasks","intervention_or_exposure":"RAG","comparator":"not extracted","endpoint":"not extracted","effect":"not extracted","risk_of_bias":"not appraised in public sidecar","directness":"primary","excerpt":"The results show that the proposed MAF-RAG significantly outperforms the baseline system, achieving a mean F1-score of 0.556, an improvement of 18.8% over the Enhanced Baseline (mean F1-score = 0.469) and a 70.0% improvement over the Legacy Baseline (mean F1-score = 0.327).","source_id":"source_5","support_kind":"candidate_source_row"}]}]}