Official evaluation mechanics

How Horizon Europe proposals are actually evaluated

Worth understanding before trusting any pre-submission score, including ours. The rules below are the ones that decide outcomes, and most of them are invisible in the call text itself.

The process

Three phases, and only the last one reaches you

01

Individual evaluation

Several independent experts read the proposal separately and each records scores and comments without seeing the others. Independence is the point: it is what makes the later disagreement informative rather than an artefact of who spoke first.

02Where scores move

Consensus discussion

The experts convene, compare their individual assessments and negotiate a single agreed position per criterion. Scores commonly move here, and a comment that survives this stage is one at least two readers were willing to defend.

03

Panel review and the ESR

A panel checks consistency across the whole set of proposals, resolves ties and ranks. The applicant eventually receives the Evaluation Summary Report — the consensus comments and scores, and nothing about how they were reached.

Bidshark reproduces this shape deliberately — independent readings, then a reconciliation step where disagreement is tested, then one senior decision — because a single pass produces one opinion with nothing to check it against.

The scale

Three criteria, five points each, half marks throughout

Each criterion is scored 0 to 5 in steps of 0.5. For most RIA and IA calls a criterion must reach 3.0 and the total must reach 10.0 of 15 — but clearing the threshold and being funded are different questions.

Excellence

0–5.0

Clarity and pertinence of the objectives, the extent to which the work is ambitious and beyond the state of the art, and the soundness of the methodology.

Impact

0–5.0

The credibility of the pathway from results to the outcomes the call expects, and the quality of the measures to maximise it — dissemination, exploitation and communication.

Implementation

0–5.0

The quality and effectiveness of the work plan, the appropriateness of effort and resources, and the capacity of the participants to deliver it.

Threshold is the floor, not the target

In oversubscribed calls, proposals well above threshold are routinely unfunded because the ranked list runs out of budget first. A total in the low tens clears the formal bar and still loses. This is why our report separates threshold status from competitive position rather than reporting a single verdict.

The rules that decide outcomes

What the call text does not tell you

Sub-criteria differ by type of action

RIA and IA, CSA, COFUND, MSCA and ERC each carry distinct sub-aspect lists. Applying one fixed Excellence/Impact/Implementation checklist to every call is structurally wrong for a meaningful share of proposals — and is the most common way a generic AI review goes astray.

Three severities, not a free-form guess

Evaluators classify what they find as a minor shortcoming, a shortcoming, or a significant weakness — the last meaning treatment so limited or ineffective that it pushes the criterion below threshold. The number follows from that classification, not the other way round.

Judged as submitted

The proposal is assessed as it stands. An evaluator may not credit it for a fix the applicant could obviously make, and the ESR carries no recommendations.

No double penalisation

One weakness may not be charged against two criteria. A thin work plan is an Implementation problem; it does not also reduce Excellence.

Hidden checks that lower Excellence on their own

Open science practices, the gender dimension in research content, SSH integration where the topic flags it, and the technical robustness of any AI involved. Each is easy to omit and each can cost points by itself.