← All notes

Horizon Europe Evaluation Threshold Explained

A proposal can clear every formal check, be technically credible and still fail on the Horizon Europe evaluation threshold. That is not a contradiction. The threshold is a minimum condition for consideration, not evidence that the proposal is competitive against the applications competing for the same budget.

For a consortium that has already committed months of technical writing, partner coordination and internal approvals, this distinction matters. Submission fixes the document. No post-submission explanation can repair an impact pathway that remains asserted rather than evidenced, a work package that does not support its claimed outcomes, or a consortium role that looks ornamental when read against the implementation criterion.

The Horizon Europe evaluation threshold is a floor

In many Horizon Europe calls, proposals are assessed against Excellence, Impact and Quality and Efficiency of Implementation. The commonly used threshold is 3 out of 5 for each criterion and 10 out of 15 overall. A proposal with 4 for Excellence, 4 for Impact and 2.5 for Implementation may therefore have a respectable-looking total but fail because it misses an individual criterion threshold.

Do not treat that pattern as universal. The definitive thresholds, weighting, award criteria and sub-criteria are those in the call conditions and the applicable evaluation form. They can differ by type of action, topic and work programme version. Two-stage calls and calls using particular procedures may also apply different arrangements. Check the documents for the exact call rather than relying on a threshold remembered from a previous submission.

A threshold is also separate from selection for funding. A proposal that reaches all thresholds enters the pool of eligible-for-funding proposals, subject to ranking and the available budget. If the budget line funds only a small number of applications, a score just above the minimum remains commercially exposed.

Admissibility, eligibility and award scores are different gates

Before evaluators assess the award criteria, the proposal must meet admissibility and eligibility requirements. These can include submission by the deadline, use of the required application parts, completeness, page limits, readability requirements and topic-specific participation conditions. A high-quality narrative does not cure a breach of these rules.

Award criteria answer a different question: how good is the eligible proposal against the published marking scheme? The evaluation summary report records this assessment, normally through criterion scores and comments. It is a mistake to fold all compliance work into a single final proofread. Page limits, administrative conditions and mandatory elements need a separate control process from the evidence needed to earn points.

How the score is actually assembled

The official 0-5 scale is not a measure of how much effort the consortium has invested. It is an evaluator judgement on whether the proposal meets the criterion, and how convincingly it does so. The published descriptors distinguish failure or non-assessability at the bottom of the scale from poor, fair, good, very good and excellent performance. Half-point scoring may be used where the evaluation form permits it.

The score follows the criterion, but the reasoning follows its sub-criteria. Under Excellence, an evaluator may test whether objectives are clear and pertinent, whether the methodology is credible, and whether interdisciplinarity, open science practices or the social sciences and humanities are handled where relevant. Under Impact, the reader tests the credibility of the pathway from project results to expected outcomes and wider impacts, alongside dissemination, exploitation and communication measures. Implementation asks whether the work plan, resources, risks, governance and consortium capacity can deliver what the application promises.

For Innovation Actions, Impact may carry a higher weighting, often 1.5. That can make a relatively modest weakness in the impact case more damaging to the final ranking than teams expect. But weighting does not remove individual thresholds. A weighted total cannot necessarily compensate for a criterion score below the required minimum.

The practical consequence is uncomfortable: a proposal does not receive points for topics it mentions. It receives a score for the quality, specificity and internal consistency of the case made under each criterion.

Where proposals lose points before they reach the threshold

Excellence: plausible is not yet substantiated

Excellence often loses points when the objectives sound ambitious but cannot be tested, measured or connected to the proposed method. A statement that the project will develop a breakthrough platform is not an objective unless the proposal states what will be developed, against what performance baseline, by when, and through which technical approach.

Method sections also fail when they substitute a sequence of activities for a rationale. Evaluators look for the chain between the research question, the method, the validation design and the expected result. If a key assumption is unsupported, or an alternative method has not been considered where risk is evident, the weakness belongs in the score, not merely in a comment.

Impact: intention is not a pathway

Impact is frequently the decisive criterion because it is where internal optimism is most visible. Market size, policy relevance and broad stakeholder lists are not an impact pathway. The proposal needs to show who will use the results, what changes for them, what barriers stand in the way, who owns the route to uptake and how progress will be evidenced during and after the project.

A KPI is useful only when it measures a step that matters. Counting workshops or website visits may support communication reporting, but it rarely demonstrates exploitation, adoption or contribution to the destination’s expected outcomes. An evaluator will also compare the claimed impacts with the work plan. If no work package funds the standardisation activity, demonstration, regulatory work or user validation needed for uptake, the claim is unsupported.

Implementation: detail must add up

Implementation scores fall when the tables contradict the narrative. Typical faults include task leaders without relevant capacity, risks that name no trigger or mitigation owner, milestones that duplicate deliverables, and person-month allocations that do not match the technical burden.

Consortium quality is not established by assembling reputable organisations. Each partner needs a necessary and credible role, with its capacity evidenced in the relevant task. A large consortium can score lower than a focused one if interfaces, decision rights and dependencies are vague. Conversely, a smaller consortium needs to show that it has not omitted capabilities essential to delivery.

Passing the threshold is not the same as being fundable

A 10 out of 15 may be enough to pass a standard overall threshold, but it may not be near the score required for funding in a competitive topic. Ranking is affected by the quality of other eligible proposals, available budget and the call’s published rules for priority and tied scores. There is no universal score that guarantees funding.

This is why a pre-submission review should not ask only, “Will this pass?” The better question is, “Which criterion contains the finding most likely to lower our rank?” A significant weakness in one section can be more valuable to identify than ten minor edits to language and layout.

A defensible check before submission

A useful review starts with the correct evaluation form for the action and work programme version. Then it should test the proposal in the same order an evaluator will: criterion first, sub-criterion second, and supporting passage third.

BidShark applies this logic through independent readings of Excellence, Impact, Implementation, consortium capacity, unsupported claims and sector-specific issues. The resulting scores are assessed against the relevant official 0-5 marking scheme, and disagreement between readers is shown and adjudicated rather than averaged away. This is automated scoring, not a claim that a human panel has read the proposal. The separate Expert Q&A package is the point at which a practising EU evaluator reads the report and answers written questions.

Whatever review method you use, require traceability. Every material criticism should identify the passage it rests on and the criterion it affects. Otherwise, the team cannot distinguish a genuine evaluation risk from a reader’s preference.

The final days before submission are not the time to make a good proposal sound more confident. They are the time to remove the unsupported claim, the uncosted dependency and the criterion-level weakness that gives a real evaluator a reason to score down.