Prompt Engineering with CoVe
Production-grade prompt engineering with Chain-of-Verification (CoVe) for factual accuracy. Combines structured outputs, few-shot learning, chain-of-thought, and verification loops to maximize LLM reliability.
Master production prompt patterns with built-in verification for factual accuracy and self-correction.
Prompt Engineering with Chain-of-Verification
Master production prompt patterns with built-in verification for factual accuracy and self-correction.
When to Use This Skill
- Building LLM apps requiring factual accuracy
- Reducing hallucinations in production systems
- Implementing self-correcting prompt chains
- Designing multi-step reasoning with verification
- Creating reliable structured outputs
- Building RAG systems with answer validation
Core Architecture: Chain-of-Verification (CoVe)
CoVe reduces hallucinations through a 4-stage pipeline:
┌─────────────┐ ┌─────────────────┐ ┌─────────────────┐ ┌──────────────┐
│ Baseline │───▶│ Plan │───▶│ Execute │───▶│ Refined │
│ Response │ │ Verification │ │ Verification │ │ Response │
└─────────────┘ └─────────────────┘ └─────────────────┘ └──────────────┘
Draft Generate Answer each Cross-check
answer verification question and refine
questions independently
Stage 1: Generate Baseline Response
BASELINE_PROMPT = """Answer the question concisely. Question: {question} Answer:"""
Stage 2: Plan Verification Questions
Generate questions that probe each factual claim:
VERIFICATION_PLANNING_PROMPT = """Given this question and answer, generate verification questions to check factual accuracy. Question: {question} Answer: {baseline_response} For each factual claim in the answer, create a verification question. Example: - If answer mentions "X was born in Y" → "Was X born in Y?" - If answer mentions "X invented Y in Z" → "Did X invent Y?" and "Was Y invented in Z?" Verification Questions:"""
Stage 3: Execute Verification
Answer each verification question independently (optionally with tools):
EXECUTE_VERIFICATION_PROMPT = """Answer this verification question accurately. Question: {verification_question} Answer (yes/no with brief explanation):"""
Stage 4: Refine Based on Verification
FINAL_REFINEMENT_PROMPT = """Refine the original answer based on verification results. Original Question: {question} Baseline Answer: {baseline_response} Verification Results: {verification_qa_pairs} Instructions: - Keep claims that passed verification - Remove or correct claims that failed - Do not add new unverified information Refined Answer:"""
Complete CoVe Implementation
from anthropic import Anthropic from typing import List, Dict import json class ChainOfVerification: def __init__(self, model="claude-sonnet-4-5"): self.client = Anthropic() self.model = model def generate(self, question: str, use_tools: bool = False) -> Dict: # Stage 1: Baseline baseline = self._get_baseline(question) # Stage 2: Plan verification v_questions = self._plan_verification(question, baseline) # Stage 3: Execute verification v_answers = self._execute_verification(v_questions, use_tools) # Stage 4: Refine final = self._refine(question, baseline, v_questions, v_answers) return { "baseline": baseline, "verification_questions": v_questions, "verification_answers": v_answers, "final_answer": final } def _get_baseline(self, question: str) -> str: response = self.client.messages.create( model=self.model, max_tokens=500, messages=[{ "role": "user", "content": f"Answer concisely:\n\n{question}" }] ) return response.content[0].text def _plan_verification(self, question: str, baseline: str) -> List[str]: prompt = f"""Generate verification questions for each factual claim. Question: {question} Answer: {baseline} Output as JSON array of strings: ["question1", "question2", ...]""" response = self.client.messages.create( model=self.model, max_tokens=500, messages=[{"role": "user", "content": prompt}] ) return json.loads(response.content[0].text) def _execute_verification( self, questions: List[str], use_tools: bool ) -> List[Dict]: results = [] for q in questions: if use_tools: answer = self._verify_with_search(q) else: answer = self._verify_self(q) results.append({"question": q, "answer": answer}) return results def _verify_self(self, question: str) -> str: response = self.client.messages.create( model=self.model, max_tokens=200, messages=[{ "role": "user", "content": f"Answer yes or no with brief explanation:\n{question}" }] ) return response.content[0].text def _refine( self, question: str, baseline: str, v_questions: List[str], v_answers: List[Dict] ) -> str: v_pairs = "\n".join( f"Q: {a['question']}\nA: {a['answer']}" for a in v_answers ) prompt = f"""Refine the answer based on verification. Original Question: {question} Baseline Answer: {baseline} Verification Results: {v_pairs} Keep verified claims. Remove/correct failed claims. Refined Answer:""" response = self.client.messages.create( model=self.model, max_tokens=500, messages=[{"role": "user", "content": prompt}] ) return response.content[0].text
CoVe Question Type Chains
Different question types need different verification strategies:
Wiki/List Questions
WIKI_TEMPLATE_PROMPT = """Create a verification question template. Example Question: Who are movie actors born in Boston? Example Template: Was [actor name] born in Boston? Actual Question: {question} Template:""" WIKI_VERIFY_PROMPT = """Using this template and baseline, generate verification questions. Question: {question} Baseline: {baseline} Template: {template} Verification Questions (one per entity):"""
Multi-Span Questions
MULTI_VERIFY_PROMPT = """Generate verification questions for each independent claim. Question: {question} Baseline: {baseline} Each claim should have its own verification question. Verification Questions:"""
Long-Form Questions
LONG_VERIFY_PROMPT = """Generate verification questions for key factual claims. Question: {question} Baseline: {baseline} Focus on verifiable facts (names, dates, numbers, relationships). Skip opinions and general statements. Verification Questions:"""
Routing to the Right Chain
ROUTER_PROMPT = """Classify this question type. Categories: - WIKI: Asks for a list of entities (people, places, things) - MULTI: Contains multiple independent sub-questions - LONG: Requires detailed explanation Question: {question} Output JSON: {{"category": "WIKI|MULTI|LONG"}}"""
Integration Patterns
CoVe + RAG: re-retrieve per verification question against the same vector store. CoVe + Tool Use: route verification questions through a web-search / SQL / API tool. CoVe + Structured Output: enforce Pydantic schema on the refined answer for downstream reliability.
Chain-of-Thought + CoVe
COT_COVE_PROMPT = """Solve this problem step by step, then verify each step. Problem: {problem} REASONING: Step 1: [reasoning] Verification: [check this step] ... FINAL ANSWER: [answer] CONFIDENCE: [high/medium/low based on verification results]"""
Best Practices
- Always verify factual claims
- Use tools when available — external verification beats self-verification
- Route by question type
- Structure outputs with Pydantic / JSON schema
- Cache system prompts and verification patterns aggressively
- Monitor verification rates and failure patterns
- Fail gracefully — when verification fails, surface the uncertainty
Common Pitfalls
- Skipping verification ("it's probably right" → hallucinations)
- Self-verification only (models confidently verify their own mistakes)
- Over-verification (not every sentence needs a question)
- Wrong chain for question type
- No structured output (parsing failures in production)
Resources
references/cove-chains.md— full chain implementations (Wiki / Multi / Long) with example tracesreferences/prompts.md— complete prompt library with rationale per prompt + anti-patternsscripts/cove-runner.py— Python CLI tool for testing CoVe on batch question sets
Source
community
Author
Tech Horizon Academy
Version
1.0
Complexity
Compatible With
Tags
