AI Is Allowed to Write Award Entries but Not to Judge Them. That Asymmetry Is Your Problem | AI award judging

The rules on AI award judging are settling into a shape almost nobody planned, and it is lopsided in a way that lands squarely on programme operators.

Applicants may use AI, subject to disclosure and an originality expectation. Reviewers may not use it at all. Look at the largest research funder in the world and the asymmetry is explicit.

On this page

The asymmetry, as it actually stands

Applicant side

Permitted

with disclosure and an originality standard

  • No outright prohibition on using AI to prepare an application
  • Work “substantially developed by AI” may fail an originality test
  • “Substantially” is deliberately undefined — no word count, no percentage
  • Disclosure-based frameworks are the direction of travel

Reviewer side

Prohibited

on confidentiality grounds, not quality grounds

  • Reviewers may not use generative AI to analyse applications or draft critiques
  • Uploading application content to an external tool breaches confidentiality
  • Reviewers sign agreements certifying this
  • Other major funders have declined to authorise AI inside review

The reviewer-side prohibition is set out plainly in the NIH notice on generative AI in peer review, and the reasoning is worth reading carefully because it is not what most people assume. The objection is not that AI judges badly. It is that AI tools offer no guarantee of where data goes, is stored, or is later used — so putting a confidential application into one is a disclosure, regardless of how good the resulting critique might be.

That distinction matters for anyone running an award programme, because it means the prohibition does not soften as models improve. A better model is not a less confidential one.

The volume problem

Follow the asymmetry through to its operational consequence and it is uncomfortable. Applicants gain a tool that lets them produce polished, well-structured, on-brief submissions at a fraction of the previous effort. Judges gain nothing.

Two effects follow, and the second is worse than the first.

Entry volume rises. The marginal cost of one more application has collapsed. Programmes should expect more entries, and specifically more entries from candidates who would previously have judged the effort not worth it.

Surface quality stops discriminating. This is the real problem. Award judging has always relied, partly unconsciously, on presentation as a proxy for substance — a well-argued, cleanly written entry signalled a well-run organisation. That proxy is now cheap. Entries that read well no longer separate themselves from entries that are good, and rubrics built around clarity of expression will reward the wrong thing.

The practical response is rubric design, not detection. AI-detection tools are unreliable enough that basing a disqualification on one is indefensible. What does work is shifting criteria toward things a language model cannot supply on an applicant’s behalf: verifiable outcomes, named references who can be contacted, specific figures that can be checked, and evidence of work that predates the application. Reweight toward what is checkable and the asymmetry stops mattering as much.

The new attack surface

Here is the part almost no award programme has considered. The governance literature around AI in grant review already flags covert prompt injection as a live concern, and award programmes are structurally more exposed than funders are.

The mechanism is simple. If any automated tool reads a submission — an AI summariser, an eligibility screener, a shortlisting assistant, a translation step — then text inside the submitted document can address that tool rather than the human reader. Instructions can be placed where a person will not see them but a parser will: rendered invisibly, buried in document metadata, or embedded in an attachment nobody opens.

Why awards are more exposed than funders. Research funders receive documents from institutions with research-integrity offices and reputations at stake. Award programmes accept open entries — frequently from commercial entrants competing for reputational advantage, sometimes with an agency preparing the submission on their behalf.

The incentive to game an automated screening step is higher, the accountability is lower, and the entry route is open to anyone.

Three defences, in order of how much they buy you. Do not place an automated tool anywhere in the path between a submission and a decision that affects it — screening and shortlisting included, since exclusion at screening is a decision. Where a tool must process submitted content, strip formatting and metadata and pass through plain text only. And treat any automated output as advisory to a named human who remains accountable for the judgement.

This belongs in the audit record too. If an automated step touched an entry, that fact should be recoverable later — part of the same defensible trail described in our guide to the award judging process.

Where we think this lands

We build award software, so treat what follows as a position rather than a neutral survey.

The confidentiality argument against AI in review looks durable to us, and it is the argument programmes should adopt rather than the quality argument. Quality arguments expire — every improvement in model capability weakens them, and a programme that told its judges “AI is not good enough” will be relitigating that annually. Confidentiality does not expire. An entrant’s submission was given to your programme for assessment by your panel, and routing it through a third-party service is a change in who holds it, whatever the service is capable of.

That framing also travels well in this region. Under the PDPO, use of personal data is bound to the purpose stated at collection, and processing submissions through an external tool your entrants were never told about sits badly against DPP1 and DPP3 — as covered in our PDPO compliance guide.

Where we are less certain: whether a blanket prohibition survives contact with judge behaviour. Judges are volunteers or lightly paid experts reading long submissions in evenings. Some will paste an entry into a chatbot to summarise it, and a rule they have signed will not reliably stop them. A prohibition nobody can observe is a policy in name only, and we do not have a good answer to that beyond making the tools inside the platform good enough that the temptation is smaller.

What to do before your next cycle

  • Write an AI clause into your entry terms — disclosure expected, originality required, no detection-based disqualification promised
  • Write the corresponding clause into your judge agreement, and give the confidentiality reason rather than a quality reason
  • Reweight your rubric toward verifiable, checkable claims and away from quality of expression
  • Audit whether any automated step already sits between submission and decision, including translation and screening
  • Plan judging capacity for higher entry volume, which may mean more panels — and therefore normalisation you did not previously need
  • Ask vendors what their platform sends to third-party AI services, and get the answer in writing — one for the RFP

The last item is the one we would push hardest, and it cuts against us as much as anyone. “Powered by AI” has become a feature claim in this category without a corresponding disclosure about where submitted data goes. If a platform summarises entries for judges, something is processing those entries. Ask what, and where.


Rethinking your entry terms or judging policy for the next cycle? Book a live demo or see pricing.

Leave a Reply

Scroll to Top

Discover more from AwardScience | Data-Driven Award & Grant Management Software

Subscribe now to keep reading and get access to the full archive.

Continue reading