Prompt Engineering for Developers: Few-Shot, Chain-of-Thought and Constraint Prompts for Code

Key takeaways

Vague requests like "make a login page" produce generic, oversized code. The post compares bad and good prompts, walks through prompts for features, bug fixes, refactoring, reviews and tests, and shows how prompting differs between a chat window, Cursor and Copilot.

Introduction

Ask an AI assistant to “create a login” and you get generic code in whatever language and framework it guesses, often far bigger than you needed. Prompt engineering for developers is mostly about removing that guesswork: stating the tech stack, requirements, constraints and expected output format up front. This post compares bad and good prompts, covers few-shot, chain-of-thought, role and constraint patterns, gives prompts for features, bug fixes, refactoring, reviews and tests, and explains how prompting differs between a chat window, Cursor and Copilot.

Prompt Basics

Bad Prompt vs Good Prompt

Bad Prompt:

"Create a login"

Problems:

  • Language/framework unclear
  • Insufficient requirements
  • No style guide Good Prompt:
"Implement a login page with Next.js 14 App Router.
Tech Stack:
- TypeScript
- React Hook Form
- Zod (validation)
- Tailwind CSS
Requirements:
1. Email/password input
2. Real-time validation (email format, password 8+ characters)
3. Loading state display
4. Error message display
5. Dark mode support
File: app/login/page.tsx
"

The good prompt works not because of any magic phrasing but because it removes decisions the model would otherwise make at random. A language model produces the most plausible continuation of what it was given; with “Create a login”, the most plausible answer is an average of every login tutorial it has seen — often a different framework, a different validation library, and a session approach you don’t use. Every line in the good version closes one of those branches. The file path matters more than it looks: it tells the model which framework conventions apply (App Router means a Server Component by default, so the form needs 'use client'), and it bounds the size of the answer to one file.

Two cautions apply to every example in this post. Version numbers drift: “Next.js 14” was current when many of these prompts were written, and a model will happily generate code for the version you name, including APIs that have since changed — name the version you actually run. And the model does not know your codebase unless you show it: if your project already has a Button component or an apiClient wrapper, say so or paste it, otherwise you get a parallel implementation that duplicates what exists.

Prompt Structure (5W1H)

1. Who (Role)

"You are a senior full-stack developer."

2. What (What to do)

"Design and implement a RESTful API."

3. Why (Purpose)

"To build a user authentication system."

4. Where (Environment)

"In a Node.js + Express + PostgreSQL environment"

5. When (Timing)

"When a user signs up"

6. How (Method)

"Using JWT token-based authentication, hashing passwords with bcrypt"

You don’t need all six in every prompt; they are a checklist for what might be missing. In practice, “Where” (stack and versions) and “How” (the approach you have already chosen) prevent the most rework, because those are the decisions the model is least likely to guess right. “Why” is underrated: saying the purpose (“for an internal admin tool with five users” versus “for a public signup page”) changes what a sensible answer looks like — the first does not need rate limiting and CAPTCHA, the second does.


Core Patterns

Zero-shot (No Examples)

Definition: Direct request without examples

"Implement a binary search function in Python"

Pros: Fast
Cons: Lower accuracy

For a well-known, self-contained task like binary search, zero-shot is usually fine — the model has seen thousands of correct implementations. Accuracy drops when the “right” answer depends on local conventions the model cannot know: your error format, your naming, your logging. That is where the next pattern helps.

Few-shot (With Examples)

Definition: Provide examples to learn patterns

"Create API responses in the following format:
Example 1 (Success):
{
  "success": true,
  "data": { "id": 1, "name": "John" },
  "error": null
}
Example 2 (Failure):
{
  "success": false,
  "data": null,
  "error": { "code": "NOT_FOUND", "message": "User not found" }
}
Now create a response format for a product lookup API."

Pros: Higher accuracy
Cons: Longer prompts

Examples are the most reliable way to pin down a format, because they show rather than describe: field names, nesting, null versus omitted fields, error code style. The model will also copy things you did not intend — if both examples use "id": 1, don’t be surprised when generated examples do too. Two or three varied examples (success, failure, an edge case) generalize better than one. In a codebase, the best few-shot example is an existing file: “write ProductController following the structure of this UserController” usually beats a paragraph of style rules.

Chain-of-Thought (Step-by-Step Thinking)

Definition: Guide AI to think step by step

"Solve the following problem step by step:
Problem: Remove duplicates and sort the array [3, 1, 4, 1, 5, 9, 2, 6]
Steps:
1. Choose duplicate removal method
2. Choose sorting algorithm
3. Analyze time complexity
4. Implement code
5. Write test cases
"

Result:

AI Response:
1. Duplicate removal: Use Set (O(n))
2. Sorting: Built-in sort (O(n log n))
3. Overall time complexity: O(n log n)
4. Code:
   def remove_duplicates_and_sort(arr):
       return sorted(set(arr))
5. Test:
   assert remove_duplicates_and_sort([3,1,4,1,5]) == [1,3,4,5]

Asking for intermediate steps helps in two ways. For the model, writing out the reasoning before the code gives later tokens something to build on, which measurably improves results on multi-step problems. For you, the steps are reviewable: if step 2 picks an approach you know is wrong for your data (say, a sort that is not stable when you need stability), you can stop there instead of debugging the final code. Many current models also have a dedicated reasoning or “thinking” mode that does this internally; with those, the explicit instruction matters less, but asking for the plan before the code is still useful because it lets you correct direction cheaply. The trap is trusting the explanation too much — a fluent step-by-step argument can still end in wrong code, so the test cases in step 5 are the part to actually run.

Role Prompting (Role Assignment)

"You are a senior backend developer with 10 years of experience.
You prioritize security, performance, and scalability.
Design the following API:
- User authentication
- Permission management
- Logging
- Error handling
"

A role on its own (“you are a senior developer”) has a smaller effect than people expect; what changes the output is the second line, which states priorities. “Prioritize security” tends to produce input validation, parameterized queries, and careful error messages that a neutral prompt might skip. If you can, replace vague priorities with concrete ones: “assume hostile input; never return stack traces; log with request ids” gives the model something to act on.

Constraint Prompting (Constraints)

"Write code following these constraints:
Required:
- Use TypeScript
- Minimize external libraries
- JSDoc comments for all functions
- Include unit tests
Prohibited:
- Using any type
- Using console.log
- Using global variables
"

Constraints are where the model’s output is easiest to verify mechanically, so pair them with tools rather than trusting compliance. “No any” is enforced by strict TypeScript and @typescript-eslint/no-explicit-any; “no console.log” by an ESLint rule. Models follow explicit prohibitions reasonably well in short answers and drift in long ones, especially when the prohibited thing is the common idiom. Phrase constraints positively when possible (“use the project logger for all output”) — it tells the model what to do instead, which works better than only saying what not to do.


Use Case Prompts

Implementing New Features

Prompt:

"Implement an infinite scroll component with React + TypeScript.
Requirements:
- Use Intersection Observer API
- Display loading state
- Error handling
- Separate into custom hook
- Support generic types
Example usage:
const { data, loading, error } = useInfiniteScroll<Post>(
  '/api/posts',
  { limit: 20 }
);
File structure:
- hooks/useInfiniteScroll.ts
- components/InfiniteScrollList.tsx
"

The “example usage” block is the strongest part of this prompt. Writing the call site you want before the implementation fixes the hook’s API — its name, generic parameter, arguments, and return shape — so the model designs to your interface instead of inventing one. It is essentially test-first design applied to prompting.

Bug Fixing

Prompt:

"Find and fix the bug in the following code:
[Paste code]
Error message:
TypeError: Cannot read property 'length' of undefined
Steps:
1. Analyze bug cause
2. Explain fix method
3. Provide fixed code
4. Add test cases
"

For bug fixes, the quality of the answer is capped by the evidence you provide. The error message alone invites a generic guess (“add a null check”), which silences the symptom without explaining why the value was undefined. Include the full stack trace, the input that triggers it, what you expected, and what you already ruled out. The most common failure mode I see with AI-assisted debugging is exactly that shallow fix: optional chaining added at the crash site while the real bug — an API response that changed shape, a race in state updates — stays in place. Asking for the cause before the fix (step 1) is what guards against it; if the explanation doesn’t account for why the value is missing, the fix is probably a patch.

Refactoring

Prompt:

"Refactor the following code:
[Paste code]
Improvements:
1. Function separation (Single Responsibility Principle)
2. Strengthen type safety
3. Add error handling
4. Performance optimization
5. Improve readability
Include explanations for each change.
"

Refactoring is the task where AI output is most dangerous to accept wholesale, because the goal is unchanged behavior and a plausible-looking rewrite can alter it subtly — reordering side effects, changing how null is handled, swallowing an exception that used to propagate. Five improvement goals at once also produce a diff too large to review. Have tests in place first (or ask for characterization tests before the refactor), request one kind of change per round, and add “do not change public signatures or behavior” explicitly.

Code Review

Prompt:

"Review the following code:
[Paste code]
Review perspectives:
1. Security vulnerabilities (OWASP Top 10)
2. Performance issues
3. Code smells
4. Best practice violations
5. Test coverage
Provide specific examples and improvement suggestions for each item.
"

AI review is good at local issues visible in the pasted code — missing input validation, SQL built by string concatenation, unhandled promise rejections — and weak at anything that depends on context it can’t see, such as whether a caller already validated the input or whether a query runs inside a transaction. Expect some false positives and ask it to rank findings by severity with the line numbers involved, so the important ones are not buried under style comments. Never paste secrets, customer data, or proprietary code into a tool your organization has not approved for it; that applies to every use case here but comes up most in review, where people paste whole files.

Writing Test Code

Prompt:

"Write test code for the following function:
[Function code]
Test framework: Jest + React Testing Library
Test cases:
1. Normal operation
2. Edge cases (empty array, null, undefined)
3. Error cases
4. Async handling
5. Mocking when needed
Add explanatory comments to each test.
"

Generated tests share a specific weakness: they are derived from the implementation, so they tend to assert what the code does rather than what it should do. If the function has a bug, the model often writes a test that locks the bug in. Describe the intended behavior in the prompt (“an empty array returns an empty result, not null”), and check at least one test by breaking the code on purpose to see that it fails. Also watch for tests that mock the very thing under test — a common result of the “mocking when needed” instruction.


Tool-Specific Strategies

ChatGPT / Claude (Chat)

Advantages:

  • Long conversation context
  • Complex explanations possible Strategy:
Step 1: Big picture
"Design a blog system architecture"
Step 2: Specification
"Write the Prisma schema for the User model"
Step 3: Implementation
"Implement the signup API"
Step 4: Improvement
"Add email duplicate check and password hashing"

Chat tools only know what is in the conversation, so this staged approach doubles as context management: each step’s output becomes the context for the next, and you review a small piece at a time. The weakness is that pasted code goes stale — after you edit the schema locally, the chat still reasons about the old version. Paste the current version of anything you changed before asking about it.

Cursor (Editor Integration)

Advantages:

  • Full project context
  • Multi-file editing Strategy:
Using @filename:
"@app.py @models.py Connect these two files
and add user authentication functionality"
Using @foldername:
"@src/components/ Add dark mode support
to all components"
Using @docs:
"@docs Refer to Next.js 14 App Router documentation
to implement dynamic routing"

Editor-integrated tools can read your files, which removes the copy-paste problem but introduces a different one: which files end up in context. Explicit @ references are more predictable than letting the tool search, and a narrow reference (two files) usually beats a whole folder, because a large context dilutes attention and costs more. A folder-wide instruction like “add dark mode to all components” produces a sweeping multi-file diff; review it file by file, and prefer running it on a branch so you can discard it cleanly. Agent modes that run commands and edit files autonomously amplify both the benefit and the risk — give them a clear stopping condition (“stop after tests pass; don’t modify files outside src/components”).

GitHub Copilot (Autocomplete)

Advantages:

  • Fast line-level completion
  • Comment-based generation Strategy:
// Write detailed comments
// TODO: Filter only active users from user array,
// sort by name, and return email list function
function getActiveUserEmails(users) {
  // Press Tab for AI to implement
  return users
    .filter(user => user.isActive)
    .sort((a, b) => a.name.localeCompare(b.name))
    .map(user => user.email);
}

With inline completion, the “prompt” is everything around the cursor: the comment, the function name, the parameter names, types, and nearby code in open files. A descriptive name like getActiveUserEmails plus a typed signature often works better than a long comment. One subtle point in this generated code: sort mutates the array it is called on. Here it sorts the new array returned by filter, so the caller’s users is untouched — but a completion that starts with users.sort(...) would reorder the caller’s data, and that kind of detail is easy to accept with one Tab.


Advanced Techniques

Iterative Refinement

1st: "Create a TODO app with React"
→ Basic implementation
2nd: "Add localStorage save functionality"
→ Add persistence
3rd: "Add drag and drop reordering"
→ UX improvement
4th: "Add dark mode support"
→ Add theme

Iterating works best with version control in the loop: commit after each accepted step, so a later step that breaks an earlier feature is a git diff away from being understood. Without that, a long chain of “now also add…” requests tends to accumulate regressions the model does not notice, because each answer focuses on the newest request.

Context Compression

Summarize long code:

"Summarize the core logic of the following code:
[Long code]
Summary format:
- Input: ...
- Processing: ...
- Output: ...
- Time complexity: ...
"

Multimodal Prompts

Image + Text:

[Attach UI design image]
"Implement this design with React + Tailwind CSS.
Include responsive design."

Meta Prompts

Prompts that write prompts:

"When I request 'Create a React component',
write a prompt to get better results.
Information to include:
- Tech stack
- Requirements
- Constraints
- File structure
"

Meta prompts are useful once, to build a reusable template for a recurring task; the result belongs in a snippet or a rules file rather than being regenerated each time. Multimodal prompts from screenshots get layout and colors roughly right, but they cannot see hover states, breakpoints, or accessibility requirements, so list those explicitly.


Real Examples

Example 1: REST API Design

Prompt:

"You are a senior backend developer.
Design and implement a RESTful API with the following requirements:
Project: Blog system
Tech stack: Node.js + Express + PostgreSQL + Prisma
Authentication: JWT
Entities:
1. User (id, email, password, name, createdAt)
2. Post (id, title, content, authorId, createdAt, updatedAt)
3. Comment (id, content, postId, authorId, createdAt)
API endpoints:
- POST /api/auth/register (signup)
- POST /api/auth/login (login)
- GET /api/posts (post list, pagination)
- POST /api/posts (create post, authentication required)
- GET /api/posts/:id (post detail)
- PUT /api/posts/:id (edit post, owner only)
- DELETE /api/posts/:id (delete post, owner only)
- POST /api/posts/:id/comments (create comment)
For each endpoint:
1. Prisma schema
2. Express router
3. Controller logic
4. Input validation (Zod)
5. Error handling
6. Example request/response
Please also provide file structure.
"

This prompt is thorough, and that is also its risk: eight endpoints times six artifacts is a very long answer, and long answers are where models truncate, abbreviate (“implement the rest similarly”), or lose consistency between files. A better use of it is as a specification in the first message, followed by one request per layer — the Prisma schema first, then auth, then posts — reviewing each before continuing. Notice also what the spec leaves open: how “owner only” is enforced, what pagination parameters look like, and what happens to comments when a post is deleted. The model will fill those gaps silently, so decide them up front.

Example 2: Algorithm Optimization

Prompt:

"Optimize the following code:
def find_duplicates(arr):
    duplicates = []
    for i in range(len(arr)):
        for j in range(i + 1, len(arr)):
            if arr[i] == arr[j] and arr[i] not in duplicates:
                duplicates.append(arr[i])
    return duplicates
Requirements:
1. Improve time complexity from O(n²) to O(n)
2. Space complexity analysis
3. Benchmark code before and after optimization
4. Explain each optimization technique
Test cases:
- [1, 2, 3, 4, 5] → []
- [1, 2, 2, 3, 3, 3] → [2, 3]
- [] → []
"

Supplying the test cases is what makes this prompt verifiable: you can run them against the answer instead of judging it by eye. They also pin down a detail the original code implies but does not state — duplicates are reported once each, in order of first appearance. An “optimized” answer using a set for the output would pass the O(n) requirement but might lose that order, and the second test case is what catches it.

Example 3: Complex UI Component

Prompt:

"Create an advanced data table component with React + TypeScript.
Features:
1. Sorting (multiple columns)
2. Filtering (text, date range, selection)
3. Pagination (server-side)
4. Row selection (single, multiple)
5. Column hide/show
6. CSV export
7. Responsive (card view on mobile)
Tech:
- TanStack Table (React Table v8)
- Tailwind CSS
- React Hook Form (filters)
- Zod (validation)
File structure:
- components/DataTable/index.tsx
- components/DataTable/types.ts
- components/DataTable/hooks/useDataTable.ts
- components/DataTable/utils.ts
Provide the role and code for each file.
"

Troubleshooting

AI Generates Too Much Code

Problem:

Generates 1000 lines at once → Hard to read

Solution:

"Break the code into small units and explain.
Step 1: Type definitions only
Step 2: Utility functions only
Step 3: Main logic only
Step 4: Test code only
"

Context Loss

Problem:

Forgets initial requirements as conversation gets long

Solution:

Request periodic summaries:
"Summarize what we've implemented so far:
- Completed features
- Technologies used
- Remaining tasks
"

Summaries help, but the more robust fix is to start a new conversation with that summary as the first message once a thread gets long. Models attend unevenly across very long contexts, and stale code from early in the conversation competes with the current version. A fresh session with a compact, current state is usually both cheaper and more accurate than continuing indefinitely.

Inconsistent Code Style

Problem:

Generates code with different styles each time

Solution:

In project root .cursorrules or system prompt:
"""
Coding Style Guide:
Language: TypeScript (strict mode)
Framework: Next.js 14 App Router
Styling: Tailwind CSS
Linter: ESLint + Prettier
Rules:
- Prefer functional programming
- Use arrow functions
- async/await (no Promise.then)
- Explicit types (no any)
- JSDoc comments required
- Handle errors with try-catch
- No magic numbers (define as constants)
Naming:
- Components: PascalCase
- Functions/variables: camelCase
- Constants: UPPER_SNAKE_CASE
- Files: kebab-case
"""

Rules files are the durable version of a constraint prompt: the tool adds them to every request, so you state conventions once. Most tools have their own mechanism — Cursor’s project rules (the newer .cursor/rules directory, with .cursorrules as the legacy single file), CLAUDE.md for Claude Code, AGENTS.md for tools that read it, and custom instructions files for Copilot. Keep them short and specific; a rules file that tries to encode an entire style guide competes with the actual task for attention, and rules that contradict your linter produce code that fails CI. Anything a formatter or linter can enforce is better enforced there, leaving the rules file for things tools cannot check, such as “use the existing apiClient for HTTP calls”.


Conclusion

Key Summary:

  1. Be Specific: Specify tech stack, requirements, constraints
  2. Step by Step: Big picture → Specification → Implementation → Improvement
  3. Provide Examples: Improve accuracy with Few-shot
  4. Manage Context: Use @filename, @foldername Effective Prompt Checklist:
- Assign role (senior developer, security expert, etc.)
- Specify tech stack (language, framework, libraries)
- List requirements (features, constraints)
- Provide examples (input/output format)
- Present file structure
- Code style guide
- Test requirements

The checklist is for spotting what a prompt is missing, not a template to fill every time. The habit that matters more than any single technique is treating output as a draft from a fast, confident colleague who has never seen your codebase: run it, read the diff, check the edge cases you care about, and keep prompts that worked so you can reuse them.

Further Reading:


Frequently Asked Questions (FAQ)

Q. Why does the AI start ignoring my constraints halfway through a long chat?

A. In long conversations, early instructions get diluted by everything that follows and can eventually fall outside the context window. Restate the key constraints when they matter, start a fresh session with a compressed summary of the decisions so far, and keep stable rules (style, libraries, forbidden patterns) in a project rules file that the tool injects on every request.