Skip to main content

    Prompt Engineering with CoVe

    Production-grade prompt engineering with Chain-of-Verification (CoVe) for factual accuracy. Combines structured outputs, few-shot learning, chain-of-thought, and verification loops to maximize LLM reliability.

    advanced
    Development & Code
    Quick Start

    Master production prompt patterns with built-in verification for factual accuracy and self-correction.

    Complete Guide

    Prompt Engineering with Chain-of-Verification

    Master production prompt patterns with built-in verification for factual accuracy and self-correction.

    When to Use This Skill

    • Building LLM apps requiring factual accuracy
    • Reducing hallucinations in production systems
    • Implementing self-correcting prompt chains
    • Designing multi-step reasoning with verification
    • Creating reliable structured outputs
    • Building RAG systems with answer validation

    Core Architecture: Chain-of-Verification (CoVe)

    CoVe reduces hallucinations through a 4-stage pipeline:

    ┌─────────────┐    ┌─────────────────┐    ┌─────────────────┐    ┌──────────────┐
    │  Baseline   │───▶│  Plan           │───▶│  Execute        │───▶│  Refined     │
    │  Response   │    │  Verification   │    │  Verification   │    │  Response    │
    └─────────────┘    └─────────────────┘    └─────────────────┘    └──────────────┘
         Draft             Generate              Answer each            Cross-check
         answer            verification          question               and refine
                           questions             independently
    

    Stage 1: Generate Baseline Response

    BASELINE_PROMPT = """Answer the question concisely. Question: {question} Answer:"""

    Stage 2: Plan Verification Questions

    Generate questions that probe each factual claim:

    VERIFICATION_PLANNING_PROMPT = """Given this question and answer, generate verification questions to check factual accuracy. Question: {question} Answer: {baseline_response} For each factual claim in the answer, create a verification question. Example: - If answer mentions "X was born in Y" → "Was X born in Y?" - If answer mentions "X invented Y in Z" → "Did X invent Y?" and "Was Y invented in Z?" Verification Questions:"""

    Stage 3: Execute Verification

    Answer each verification question independently (optionally with tools):

    EXECUTE_VERIFICATION_PROMPT = """Answer this verification question accurately. Question: {verification_question} Answer (yes/no with brief explanation):"""

    Stage 4: Refine Based on Verification

    FINAL_REFINEMENT_PROMPT = """Refine the original answer based on verification results. Original Question: {question} Baseline Answer: {baseline_response} Verification Results: {verification_qa_pairs} Instructions: - Keep claims that passed verification - Remove or correct claims that failed - Do not add new unverified information Refined Answer:"""

    Complete CoVe Implementation

    from anthropic import Anthropic from typing import List, Dict import json class ChainOfVerification: def __init__(self, model="claude-sonnet-4-5"): self.client = Anthropic() self.model = model def generate(self, question: str, use_tools: bool = False) -> Dict: # Stage 1: Baseline baseline = self._get_baseline(question) # Stage 2: Plan verification v_questions = self._plan_verification(question, baseline) # Stage 3: Execute verification v_answers = self._execute_verification(v_questions, use_tools) # Stage 4: Refine final = self._refine(question, baseline, v_questions, v_answers) return { "baseline": baseline, "verification_questions": v_questions, "verification_answers": v_answers, "final_answer": final } def _get_baseline(self, question: str) -> str: response = self.client.messages.create( model=self.model, max_tokens=500, messages=[{ "role": "user", "content": f"Answer concisely:\n\n{question}" }] ) return response.content[0].text def _plan_verification(self, question: str, baseline: str) -> List[str]: prompt = f"""Generate verification questions for each factual claim. Question: {question} Answer: {baseline} Output as JSON array of strings: ["question1", "question2", ...]""" response = self.client.messages.create( model=self.model, max_tokens=500, messages=[{"role": "user", "content": prompt}] ) return json.loads(response.content[0].text) def _execute_verification( self, questions: List[str], use_tools: bool ) -> List[Dict]: results = [] for q in questions: if use_tools: answer = self._verify_with_search(q) else: answer = self._verify_self(q) results.append({"question": q, "answer": answer}) return results def _verify_self(self, question: str) -> str: response = self.client.messages.create( model=self.model, max_tokens=200, messages=[{ "role": "user", "content": f"Answer yes or no with brief explanation:\n{question}" }] ) return response.content[0].text def _refine( self, question: str, baseline: str, v_questions: List[str], v_answers: List[Dict] ) -> str: v_pairs = "\n".join( f"Q: {a['question']}\nA: {a['answer']}" for a in v_answers ) prompt = f"""Refine the answer based on verification. Original Question: {question} Baseline Answer: {baseline} Verification Results: {v_pairs} Keep verified claims. Remove/correct failed claims. Refined Answer:""" response = self.client.messages.create( model=self.model, max_tokens=500, messages=[{"role": "user", "content": prompt}] ) return response.content[0].text

    CoVe Question Type Chains

    Different question types need different verification strategies:

    Wiki/List Questions

    WIKI_TEMPLATE_PROMPT = """Create a verification question template. Example Question: Who are movie actors born in Boston? Example Template: Was [actor name] born in Boston? Actual Question: {question} Template:""" WIKI_VERIFY_PROMPT = """Using this template and baseline, generate verification questions. Question: {question} Baseline: {baseline} Template: {template} Verification Questions (one per entity):"""

    Multi-Span Questions

    MULTI_VERIFY_PROMPT = """Generate verification questions for each independent claim. Question: {question} Baseline: {baseline} Each claim should have its own verification question. Verification Questions:"""

    Long-Form Questions

    LONG_VERIFY_PROMPT = """Generate verification questions for key factual claims. Question: {question} Baseline: {baseline} Focus on verifiable facts (names, dates, numbers, relationships). Skip opinions and general statements. Verification Questions:"""

    Routing to the Right Chain

    ROUTER_PROMPT = """Classify this question type. Categories: - WIKI: Asks for a list of entities (people, places, things) - MULTI: Contains multiple independent sub-questions - LONG: Requires detailed explanation Question: {question} Output JSON: {{"category": "WIKI|MULTI|LONG"}}"""

    Integration Patterns

    CoVe + RAG: re-retrieve per verification question against the same vector store. CoVe + Tool Use: route verification questions through a web-search / SQL / API tool. CoVe + Structured Output: enforce Pydantic schema on the refined answer for downstream reliability.

    Chain-of-Thought + CoVe

    COT_COVE_PROMPT = """Solve this problem step by step, then verify each step. Problem: {problem} REASONING: Step 1: [reasoning] Verification: [check this step] ... FINAL ANSWER: [answer] CONFIDENCE: [high/medium/low based on verification results]"""

    Best Practices

    1. Always verify factual claims
    2. Use tools when available — external verification beats self-verification
    3. Route by question type
    4. Structure outputs with Pydantic / JSON schema
    5. Cache system prompts and verification patterns aggressively
    6. Monitor verification rates and failure patterns
    7. Fail gracefully — when verification fails, surface the uncertainty

    Common Pitfalls

    • Skipping verification ("it's probably right" → hallucinations)
    • Self-verification only (models confidently verify their own mistakes)
    • Over-verification (not every sentence needs a question)
    • Wrong chain for question type
    • No structured output (parsing failures in production)

    Resources

    • references/cove-chains.md — full chain implementations (Wiki / Multi / Long) with example traces
    • references/prompts.md — complete prompt library with rationale per prompt + anti-patterns
    • scripts/cove-runner.py — Python CLI tool for testing CoVe on batch question sets
    Skill Details

    Source

    community

    Author

    Tech Horizon Academy

    Version

    1.0

    Complexity

    Compatible With

    Claude code
    Claude api
    cursor

    Tags

    chain-of-verification
    hallucination-reduction
    production-llm
    prompt-engineering
    structured-output
    Need Help?
    Learn more about using Claude Skills effectively