Research dataset · Version 1

The dissertation prototype.

The original MSc pipeline proved the core proposition: PubMed abstracts can be retrieved programmatically, converted into a consistent schema with a locally hosted LLM, cleaned deterministically, scored, and then analysed as structured evidence.

191retrieved
126retained
60.7mean evidence
45.2mean confidence
18–100score range

Evidence strength distribution

Project-defined scoring bands used to compare heterogeneous studies. These are analytical categories, not clinical grading thresholds.

Strong 80–100
22
Moderate 60–79
38
Weak 40–59
57
Insufficient <40
9

What V1 showed

Human evidence was the strongest subset; animal and in-vitro records were useful for biological plausibility but carried lower direct clinical relevance and lower extraction confidence.

Evidence subsetRecordsMean evidenceMean confidence
Human5368.862.1
Animal / preclinical4857.233.4
Mixed1752.835.3
In vitro845.125.6
Scoring architecture

A transparent 100-point framework.

Evidence and confidence are separated so that a well-designed study with an incomplete abstract is not treated the same as a clearly extracted but methodologically weak study.

Study design
30
maximum points
Result + statistics
20
maximum points
Sample size
15
maximum points
Comparator
15
maximum points
Safety reporting
10
maximum points
Clinical relevance
10
maximum points
Primary V1 interpretation

Evidence maturity is uneven.

Clinically mature peptide-related therapies were more likely to be represented by human studies and stronger evidence profiles. More experimental peptides were generally supported by sparser and more preclinical evidence. The important analytical result is the structure of the evidence base, not a claim that a particular compound is effective or ineffective.