
Senior Prompt Engineer - Evaluation, Scoring and Reliability
- or -
Post a project like this29
£21/hr(approx. $28/hr)
- Posted:
- Proposals: 14
- Remote
- #4514420
- Open for Proposals
Autonomous AI Agent Developer | Full-Stack SEO Expert | Senior System Architect

AI & Automation Dev | OpenAI, n8n, Zapier, Make | PHP/Laravel, Node, Django | React, Vue | Shopify, WooCommerce, WordPress | Twilio, Airtable, Supabase | GHL
⭐⭐⭐⭐⭐ AI Automation and Workflow Engineer | Senior Web and Software Developer | Data Analyst | Data Engineer | Analytics Engineer | Business Analyst | Copywriter
We Fix Slow, Buggy Websites and Turn Traffic Into Sales | WordPress, Shopify, Custom Sites, Technical SEO/AEO/CRO

Full Stack Developer - Laravel, WordPress, Opencart, Shopify Design | SEO, Google Ads, Facebook Meta Ads Expert


132341081022304912119380111872861306289511374771367494413737660374646113496805831765506971
Description
Experience Level: Expert
We need an experienced prompt engineer to design reliable LLM-based evaluation and scoring workflows.
You must understand semantic alignment, rubric decomposition, structured outputs, calibration, prompt sensitivity, model bias, context handling, and failure analysis - not simply write persuasive instructions. The work requires improving scoring consistency across paired inputs, distinguishing genuine meaning from superficial language overlap, and testing performance against paraphrases, contradictions, omissions, formatting changes, and adversarial cases.
To apply, describe your evaluation methodology, experience building LLM judges, approach to measuring reliability, and one difficult consistency problem you solved. Include anonymised examples and tools used.
Responses to this advert generated by LLMs will be auto-rejected
To eliminate AI responses to the questions below, there is very special question below!
90% of AIs answer it incorrectly (its an unstable question)
Advanced prequalification questions
1. You are designing an LLM judge to determine whether a supplier’s reply meaningfully addresses a buyer’s request. Replies using identical wording receive higher scores than accurate paraphrases. What is the best first intervention?
* A. Increase temperature to expose the judge to more interpretations
* B. Instruct the judge to assess semantic coverage, explicitly distinguish lexical overlap from meaning, and provide contrasting scored examples
* C. Remove the buyer’s original request from the prompt
* D. Reward replies containing the buyer’s most frequent terms
2. A reply correctly addresses four requirements but confidently contradicts a fifth, critical requirement. Your judge often gives it a high overall score. Which rubric design is strongest?
* A. Ask for a holistic score based on overall helpfulness
* B. Score each requirement separately, identify contradictions, apply a critical-failure rule, and then calculate the overall result
* C. Tell the judge to penalise inaccurate responses
* D. Increase the maximum score so errors have less influence
3. Two semantically equivalent replies receive materially different scores when their order, formatting, or writing style changes. Which evaluation approach best isolates the problem?
* A. Run each version repeatedly, randomise irrelevant presentation features, and measure score variance by transformation type
* B. Reduce both replies to the same word count
* C. Set temperature to 1.0 and compare one result from each
* D. Ask a larger model which reply is better
4. An LLM judge must assess whether a response covers every material element of a complex request. Which prompt architecture is most robust?
* A. Ask for one immediate score to minimise token usage
* B. First extract atomic requirements, then map response evidence to each requirement, classify coverage, and derive a schema-constrained score
* C. Ask the model to explain its reasoning before reading the response
* D. Provide a long definition of quality and request a percentage
5. Your judge achieves 92% agreement with human labels, but performance falls sharply on short, indirect, or poorly written inputs. What should happen next?
* A. Deploy because overall agreement exceeds 90%
* B. Add more examples of well-written inputs
* C. Segment results by input characteristics, build targeted adversarial cases, revise the rubric, and recalibrate against disputed human-labelled examples
* D. Increase the system prompt’s authority and repeat the same test
6. A candidate claims that requiring detailed chain-of-thought will make scoring more accurate and auditable. Which response shows the strongest understanding?
* A. Correct; longer reasoning always improves accuracy
* B. Correct, provided temperature is zero
* C. Not necessarily; use explicit intermediate classifications and evidence fields, validate outputs, and evaluate accuracy empirically without depending on unrestricted hidden reasoning
* D. Incorrect; LLM judges should return only a number
Please answer each question with A,B,C orD.
For each, add your percentage certainty of your answer and your reasoning.
You must understand semantic alignment, rubric decomposition, structured outputs, calibration, prompt sensitivity, model bias, context handling, and failure analysis - not simply write persuasive instructions. The work requires improving scoring consistency across paired inputs, distinguishing genuine meaning from superficial language overlap, and testing performance against paraphrases, contradictions, omissions, formatting changes, and adversarial cases.
To apply, describe your evaluation methodology, experience building LLM judges, approach to measuring reliability, and one difficult consistency problem you solved. Include anonymised examples and tools used.
Responses to this advert generated by LLMs will be auto-rejected
To eliminate AI responses to the questions below, there is very special question below!
90% of AIs answer it incorrectly (its an unstable question)
Advanced prequalification questions
1. You are designing an LLM judge to determine whether a supplier’s reply meaningfully addresses a buyer’s request. Replies using identical wording receive higher scores than accurate paraphrases. What is the best first intervention?
* A. Increase temperature to expose the judge to more interpretations
* B. Instruct the judge to assess semantic coverage, explicitly distinguish lexical overlap from meaning, and provide contrasting scored examples
* C. Remove the buyer’s original request from the prompt
* D. Reward replies containing the buyer’s most frequent terms
2. A reply correctly addresses four requirements but confidently contradicts a fifth, critical requirement. Your judge often gives it a high overall score. Which rubric design is strongest?
* A. Ask for a holistic score based on overall helpfulness
* B. Score each requirement separately, identify contradictions, apply a critical-failure rule, and then calculate the overall result
* C. Tell the judge to penalise inaccurate responses
* D. Increase the maximum score so errors have less influence
3. Two semantically equivalent replies receive materially different scores when their order, formatting, or writing style changes. Which evaluation approach best isolates the problem?
* A. Run each version repeatedly, randomise irrelevant presentation features, and measure score variance by transformation type
* B. Reduce both replies to the same word count
* C. Set temperature to 1.0 and compare one result from each
* D. Ask a larger model which reply is better
4. An LLM judge must assess whether a response covers every material element of a complex request. Which prompt architecture is most robust?
* A. Ask for one immediate score to minimise token usage
* B. First extract atomic requirements, then map response evidence to each requirement, classify coverage, and derive a schema-constrained score
* C. Ask the model to explain its reasoning before reading the response
* D. Provide a long definition of quality and request a percentage
5. Your judge achieves 92% agreement with human labels, but performance falls sharply on short, indirect, or poorly written inputs. What should happen next?
* A. Deploy because overall agreement exceeds 90%
* B. Add more examples of well-written inputs
* C. Segment results by input characteristics, build targeted adversarial cases, revise the rubric, and recalibrate against disputed human-labelled examples
* D. Increase the system prompt’s authority and repeat the same test
6. A candidate claims that requiring detailed chain-of-thought will make scoring more accurate and auditable. Which response shows the strongest understanding?
* A. Correct; longer reasoning always improves accuracy
* B. Correct, provided temperature is zero
* C. Not necessarily; use explicit intermediate classifications and evidence fields, validate outputs, and evaluate accuracy empirically without depending on unrestricted hidden reasoning
* D. Incorrect; LLM judges should return only a number
Please answer each question with A,B,C orD.
For each, add your percentage certainty of your answer and your reasoning.
Jonathan F.
100% (4)Projects Completed
3
Freelancers worked with
3
Projects awarded
4%
Last project
4 Dec 2023
United Kingdom
New Proposal
Login to your account and send a proposal now to get this project.
Log inClarification Board Ask a Question
-
There are no clarification messages.
We collect cookies to enable the proper functioning and security of our website, and to enhance your experience. By clicking on 'Accept All Cookies', you consent to the use of these cookies. You can change your 'Cookies Settings' at any time. For more information, please read ourCookie Policy
Cookie Settings
Accept All Cookies