← All notes

How Are Horizon Europe Proposals Scored in Practice?

A Horizon Europe proposal can be technically excellent, assembled by credible partners and still lose at evaluation because its case is not evidenced where the evaluation form requires it. That is the practical answer to the question, how are Horizon Europe proposals scored: evaluators score what is in the submitted document against the published award criteria, not what the consortium knows, intended to include, or could explain in a meeting.

For a consortium that has already spent months on partner coordination, technical design, work packages and budget negotiations, that distinction is expensive. Once the proposal is submitted, there is normally no opportunity to repair an unclear impact pathway, a resource calculation that does not add up, or an unsupported claim of market uptake.

The evaluation starts before the score

Admissibility and eligibility are not award criteria. They are gatekeeping checks. A proposal may be excluded before scientific or technical quality is considered if it fails requirements on format, page limits, submission, completeness, topic scope, legal entity status or other conditions stated in the call.

Passing those checks does not earn points. It merely puts the proposal into evaluation.

The published call conditions, work programme and evaluation form then govern the scoring. Applicants should resist relying on a remembered template from a previous Research and Innovation Action or Innovation Action. Horizon Europe uses common principles, but the applicable form, sub-criteria, thresholds, weighting and ranking rules can vary by action type and call. The call documentation is the controlling document.

How Horizon Europe proposals are scored against three criteria

Most Horizon Europe proposals are assessed under three award criteria: Excellence, Impact, and Quality and efficiency of the implementation. Each criterion is scored on an official 0-5 scale, usually allowing half-point increments.

The score bands are deceptively simple. A score of 5 means excellent, with all relevant aspects successfully addressed and only minor shortcomings. A 4 means very good, with several relevant aspects addressed well but a small number of shortcomings. A 3 means good, but with important shortcomings. A 2 indicates that the criterion is broadly addressed but contains significant weaknesses. A 1 means inadequate, while 0 means the proposal fails to address the criterion or cannot be assessed because information is missing or incomplete.

This is not a checklist in which enough positive statements automatically produce a 5. Evaluators make a judgement about the quality, credibility and completeness of the evidence under each criterion. A claim without a mechanism, baseline, owner, evidence source or realistic assumption may be treated as an unsubstantiated assertion rather than a strength.

Excellence: is the proposed work convincing?

Under Excellence, evaluators examine the clarity and pertinence of the objectives, the ambition beyond the state of the art, and the soundness of the methodology. Depending on the action and topic, this can include interdisciplinarity, social sciences and humanities integration, sex and gender analysis, and open science practices.

A common scoring problem is confusing technical detail with methodological credibility. Ten pages describing an architecture, a laboratory process or an AI model do not necessarily explain why the approach will test the key hypotheses, manage uncertainty or produce the claimed advance.

Evaluators look for a logical chain: a defined problem, precise objectives, an approach capable of meeting them, and measurable evidence that the approach has succeeded. If the objectives are broad but the work plan only delivers exploratory activity, the gap is visible. If the claimed advance depends on access to data, facilities or standards not secured by the consortium, that dependency needs treatment, not optimism.

Impact: can the proposal credibly produce the outcomes?

Impact is often where ambitious proposals overstate rather than demonstrate. Evaluators assess the credibility of the pathways towards the expected outcomes and impacts in the destination, alongside the scale and significance of the contribution. They also examine measures to maximise impact, including dissemination, exploitation and communication.

The distinction matters. An expected outcome is not a slogan to repeat from the topic text. It is a condition the project must plausibly help create. The proposal needs to show how its outputs will be taken up, by whom, under what incentives or constraints, and on what timescale.

A credible pathway will connect project outputs to users, intermediaries, regulatory conditions, investment requirements, standards, procurement, business models or public-service adoption as appropriate. The relevant route depends on the topic. A research infrastructure project, a clinical research proposal and a deep-tech Innovation Action do not carry the same commercialisation burden. But all must show more than a final demonstration and a list of dissemination channels.

Weaknesses frequently appear in the impact section as uncosted exploitation plans, vague stakeholder engagement, KPIs with no baseline, or global market figures that do not establish a route to adoption. Statements such as “the consortium’s networks will ensure uptake” invite scrutiny. Which partners, which users, what agreement, what activity, and what evidence would show uptake?

Quality and efficiency of implementation: can this consortium deliver?

Implementation covers the quality and effectiveness of the work plan, the assessment and management of risks, the appropriateness of effort and resource allocation, and the capacity and role of each participant. The consortium is assessed here as part of delivery, not as a collection of impressive biographies.

Evaluators read across Part B. They compare work-package effort with tasks and deliverables; compare governance with decision-making risks; compare the critical path with dependencies; and compare partner expertise with assigned responsibility. Internal inconsistency is damaging because it makes the plan look untested.

A work package is not credible merely because it has objectives, tasks and deliverables. It must have enough person-months for the work described, a named lead with relevant capacity, decision points that occur before dependent tasks, and risks that can be acted on. A risk register that says “mitigate through close monitoring” without a trigger, owner or contingency will rarely reassure an evaluator.

Thresholds, weighting and the ranking problem

For many standard Horizon Europe actions, each criterion must meet a threshold of 3 out of 5, with an overall threshold of 10 out of 15. A proposal scoring 4.5 for Excellence, 2.5 for Impact and 4.5 for Implementation may therefore fail despite a strong overall average. One weak criterion can be decisive.

For Innovation Actions, the Impact score is commonly weighted by 1.5 for ranking purposes. This does not mean a weak Impact section can be offset by Excellence. The individual criterion threshold still applies where the call uses the standard rule. It means that, among proposals above threshold, a difference in Impact can carry more ranking consequence.

Calls can specify different thresholds, weighting or rules for dealing with ex aequo proposals. Some use additional conditions or adapted evaluation forms. Check the call-specific annex rather than applying a generic 10-out-of-15 rule without verification.

Funding is not awarded simply because a proposal passes threshold. The real contest is ranking against other eligible, admissible proposals within the available budget. A score that would have been fundable in one call may sit below the funding line in another. That is why “good enough” is not a serious internal standard for a highly subscribed topic.

What the evaluation process does with disagreement

Evaluators initially assess proposals individually against the evaluator brief and official form. Their readings can differ, especially where a claim is ambiguous or the evidence is dispersed across sections. A consensus process reconciles those views into a consensus assessment and scores, followed by further panel or ranking steps where required.

Applicants should not assume that disagreement is random. It often exposes a drafting problem. If one evaluator can infer a credible delivery mechanism while another cannot locate it, the proposal has made its case dependent on favourable interpretation. The real panel is not obliged to supply that interpretation.

The Evaluation Summary Report records the final outcome and the key strengths and weaknesses identified during evaluation. It is not a full commentary on every sentence. Absence of a criticism does not prove that a section was strong; it may simply not have been decisive in consensus.

A pre-submission check should reproduce the pressure points

Before submission, review the proposal criterion by criterion using the exact form for the call. Test every major claim against the text that supports it. Can an evaluator identify the evidence without searching across the document? Does each expected outcome have a pathway, assumptions, ownership and indicators? Do the work plan, Gantt chart, tables, effort figures and risk register tell the same story?

This is where an independent reading has commercial value. The drafting team has usually normalised its own assumptions by the final week. A controlled pre-submission assessment, such as BidShark’s, can identify where specialist readings diverge, quote the passages behind the finding and force a choice about what to repair. It is an automated assessment against the applicable form, not a human panel verdict; where a real evaluator’s judgement is required, that should be sought explicitly.

You only get one submitted version. Treat every unsupported assertion, threshold risk and inconsistency as something the real evaluators may see first - because they will.