Skip to content

How to Choose Prompting, RAG, or Fine-Tuning for an AI App

Learn when to use prompting, RAG, fine-tuning, or reinforcement fine-tuning for an AI application. Build a baseline, measure failures, and choose the right customization path.

How to Choose Prompting, RAG, or Fine-Tuning for an AI App

On this page

When an AI application produces weak results, developers often assume the model needs fine-tuning. That is frequently the wrong first step. AWS's latest guidance for generative AI customization lays out a progression that starts with an existing model and moves through prompt engineering, Retrieval-Augmented Generation (RAG), fine-tuning, reinforcement fine-tuning, and more specialized customization only when simpler approaches stop meeting the application's requirements.

This tutorial turns that progression into a practical decision process. You will build a simple baseline, test whether better prompting is enough, add external knowledge with RAG, and identify the point where fine-tuning becomes technically justified. The examples use Amazon Bedrock terminology, but the underlying decision process applies to most modern AI application stacks.

Start With the Problem, Not the Model

Before changing a model, define exactly what is going wrong. An AI assistant may produce incorrect answers because it lacks current information, uses the wrong response format, does not understand domain-specific language, or fails to follow a particular behavior consistently. Those problems require different solutions. AWS's current customization guidance explicitly recommends starting with the simplest approach and increasing customization only when the current level cannot satisfy the use case.

For example, imagine a support assistant that must answer questions about your company's products. If the problem is that the assistant does not know a product released last week, fine-tuning is unlikely to be the first solution. The missing information belongs in a knowledge source that the application can retrieve at request time. If the problem is instead that the model consistently needs to produce responses in a highly specific format, prompt design or supervised fine-tuning may be more appropriate.

Build a Baseline With a Normal Model Call

Start by measuring what a foundation model can already do without customization. A foundation model is a general-purpose model that has already been trained before your application uses it. Amazon Bedrock provides access to multiple foundation-model providers through its managed service.

Your first experiment should use a small set of representative prompts rather than a single example. Record the model's response, the expected response, the failure type, and any measurable quality criteria. This baseline becomes important later because customization should be justified by measurable improvement rather than by the assumption that a more complicated pipeline must be better.

from openai import OpenAI

client = OpenAI()

response = client.responses.create(
    model="your-selected-model",
    input=(
        "Answer this customer question using a concise "
        "professional tone: How can I reset my account password?"
    )
)

print(response.output_text)

The exact model and SDK can vary by provider. The important part of this baseline is that you establish the application's current behavior before introducing retrieval or training.

Try Prompt Engineering First

Prompt engineering means changing the instructions and input structure given to the model without changing the model's learned parameters. It is usually the lowest-effort customization because no training dataset or model-training job is required. AWS places direct use of foundation models and prompt engineering at the simplest end of its customization spectrum.

Instead of asking a support model to β€œanswer the customer,” define the role, output requirements, limitations, and desired behavior explicitly:

system_prompt = """
You are a customer support assistant.

Rules:
1. Answer in no more than 120 words.
2. Use only information supplied in the conversation.
3. If the required information is missing, say so.
4. Do not invent product policies.
5. Give troubleshooting steps as a numbered list.
"""

Test this revised prompt against the same evaluation set used for your baseline. If the quality improves enough to meet the application's requirements, stop there. Adding retrieval or fine-tuning to a problem already solved by clearer instructions increases complexity without necessarily providing a useful benefit.

Recognize When Prompting Is Not Enough

Prompt engineering cannot supply information that the model does not have access to. If your application depends on private company documents, changing product specifications, internal procedures, or frequently updated information, putting more instructions into the prompt is not the same as giving the model a reliable knowledge source.

This is the point where Retrieval-Augmented Generation becomes useful. RAG retrieves relevant information from an external knowledge base at inference time and includes that information in the model's context. The model does not need to memorize the entire knowledge base during training because the application supplies relevant material when the user asks a question.

Add RAG for Private or Changing Knowledge

A typical RAG pipeline has four stages: ingest documents, split them into searchable pieces, create representations that can be retrieved, and supply the most relevant pieces to the model with the user's question. This makes RAG particularly useful when information changes more frequently than model training should occur.

Consider a company that maintains hundreds of product manuals. Instead of fine-tuning a model every time a manual changes, the application can index the manuals and retrieve the relevant sections when a customer asks a question. This separates the model's general language ability from the application's changing factual knowledge.

def answer_question(question, retrieved_documents):
    context = "\n\n".join(retrieved_documents)

    prompt = f"""
Answer the user's question using the supplied documents.

Documents:
{context}

User question:
{question}

If the documents do not contain enough information,
say that the available documentation is insufficient.
"""

    return generate_with_model(prompt)

The critical part is not the retrieval library itself. It is whether the retrieved documents actually contain the information needed to answer the question. A sophisticated retrieval system with poor source documents can still produce poor answers.

Measure Retrieval Before Blaming the Model

RAG introduces a new failure point: retrieval. A model may generate a perfectly reasonable answer from the context it receives while still being wrong for the user's actual question because the retrieval layer supplied irrelevant documents.

Test retrieval independently. For each evaluation question, record whether the correct document or passage appeared among the retrieved results. This separates retrieval failures from generation failures. If the correct information is consistently retrieved but the final answer is poor, improving the model prompt or generation configuration may be more useful than rebuilding the search system.

Observed failure Likely area to investigate
Correct information exists but is not retrieved Retrieval, indexing, chunking, or search configuration
Correct information is retrieved but ignored Prompt structure or model behavior
Model follows the information but uses the wrong style Prompt engineering or fine-tuning
Model lacks a specialized behavior across many examples Fine-tuning evaluation
Success can be measured with a reliable reward Reinforcement fine-tuning evaluation

Know When Fine-Tuning Makes Sense

Fine-tuning changes model parameters using task-specific training examples. Amazon Bedrock currently supports fine-tuning for a range of foundation models, including selected Amazon Nova and Meta Llama models, with availability varying by model and AWS Region.

Fine-tuning becomes more interesting when the application needs a behavior that is difficult to establish reliably through instructions alone. Examples include consistent output formatting, specialized classification behavior, domain-specific response patterns, or a particular style that must be reproduced across many inputs.

It is generally not the right tool for information that changes constantly. If a product catalog changes every day, training the model on today's catalog does not automatically make tomorrow's catalog available. A retrieval system can provide current information without repeatedly modifying model parameters.

Prepare a Fine-Tuning Dataset

Amazon Bedrock's current fine-tuning workflow uses training data stored in Amazon Simple Storage Service (Amazon S3), with the required record format depending on the selected model and customization method. AWS documentation describes JSON Lines, or JSONL, files where each line represents a training record.

A simple example can look like this:

{"prompt":"Classify this ticket: I cannot sign in after changing my phone number.","completion":"account_access"}

{"prompt":"Classify this ticket: The package arrived with a damaged screen.","completion":"product_damage"}

{"prompt":"Classify this ticket: I want to know when my order will arrive.","completion":"shipping_status"}

The quality of these examples matters more than simply producing a large file. Each example should demonstrate the behavior you actually want the customized model to learn. Inconsistent labels, ambiguous instructions, and contradictory answers can teach the model the wrong pattern.

Create a Validation Set

Do not use every example for training. Keep a separate validation set so you can measure how the customized model behaves on examples that were not used to update its parameters. AWS's fine-tuning workflow supports validation data where supported by the selected model and customization method.

Your evaluation set should contain realistic edge cases rather than only easy examples. For a customer-support classifier, include short messages, spelling mistakes, overlapping categories, incomplete questions, and cases where the correct answer is uncertain. The objective is to find out whether the model learned the intended behavior rather than whether it memorized the training examples.

Submit a Fine-Tuning Job

After preparing the data and permissions, Amazon Bedrock allows you to create a fine-tuning job through its console or API. The API requires information such as the base model, custom model name, training data location, output location, hyperparameters, and an Identity and Access Management (IAM) service role.

request = {
    "baseModelIdentifier": "YOUR_BASE_MODEL",
    "customModelName": "support-classifier-v1",
    "jobName": "support-classifier-training",
    "hyperParameters": {
        "epochCount": "2"
    },
    "trainingDataConfig": {
        "s3Uri": "s3://your-bucket/train/"
    },
    "outputDataConfig": {
        "s3Uri": "s3://your-bucket/output/"
    },
    "roleArn": "YOUR_BEDROCK_SERVICE_ROLE_ARN"
}

The exact hyperparameters and supported configuration depend on the model you select. AWS notes that training duration depends on factors including dataset size, token count, number of epochs, and batch size.

Do Not Fine-Tune Just Because You Can

Fine-tuning introduces additional data preparation, training, evaluation, model management, and cost considerations. AWS currently charges model customization based on the amount of training data processed and the number of epochs, while customized models also incur model-storage costs.

That makes a simple experiment useful before committing to a large dataset. Compare the existing model, improved prompt, RAG pipeline, and fine-tuned model using the same evaluation set. If fine-tuning provides only a small improvement over a much simpler approach, the additional operational complexity may not be justified for your application.

Consider Reinforcement Fine-Tuning for Measurable Goals

Supervised fine-tuning teaches from example inputs and desired outputs. Reinforcement fine-tuning takes a different approach: the model receives feedback through reward signals that represent how well its responses satisfy a defined objective. AWS currently supports reinforcement fine-tuning for selected models and describes it as useful when response quality can be measured through rules, judges, execution results, or other reward functions.

This approach can be useful when there are many acceptable answers but some are clearly better than others. Code generation is one example because a generated program can be executed and graded by tests. A structured business workflow can also provide measurable signals if successful completion can be verified programmatically.

def reward(response, expected_behavior):
    score = 0

    if passes_format_check(response):
        score += 1

    if passes_accuracy_check(response):
        score += 2

    if meets_business_rule(response):
        score += 2

    return score

The reward function becomes the definition of success, so it must be designed carefully. If the reward measures something unrelated to the real application objective, the model can improve its score without actually becoming more useful.

Use a Simple Decision Tree

You can turn the customization process into a repeatable decision tree. Start with the simplest available model and determine what is missing. If the problem is instructions or formatting, improve the prompt. If the problem is missing or changing knowledge, add retrieval. If the problem is persistent specialized behavior that remains after those steps, evaluate fine-tuning. If success can be measured with a reliable reward and supervised examples are insufficient, investigate reinforcement fine-tuning. This progression follows the general escalation approach described in AWS's current customization framework.

  1. General task: Use a foundation model directly.
  2. Behavior or formatting problem: Improve the prompt.
  3. Missing or changing knowledge: Add RAG.
  4. Persistent specialized behavior: Evaluate supervised fine-tuning.
  5. Measurable reward: Consider reinforcement fine-tuning.
  6. Highly specialized requirements: Investigate deeper model customization only after the earlier approaches are tested.

Combine Techniques Instead of Treating Them as Alternatives

These techniques are not mutually exclusive. A production application can use a fine-tuned model together with RAG, for example. The customized model can learn how the application should respond, while retrieval supplies current information that should not be baked into the model parameters.

The same application can also use carefully designed system instructions on top of both. The architecture should assign each technique a clear responsibility: prompts control immediate behavior, retrieval supplies external knowledge, and model customization changes learned behavior. Keeping those responsibilities separate makes the system easier to test and maintain.

Run the Same Evaluation After Every Change

The safest way to choose a customization method is to treat every change as an experiment. Keep one evaluation set and run the baseline model against it first. Then measure the result after prompt improvements, after retrieval is added, and after any fine-tuning job completes.

Track more than one metric. Depending on the application, useful measurements can include factual accuracy, formatting compliance, retrieval accuracy, task completion rate, latency, token usage, and cost per successful request. A method that improves accuracy while multiplying cost may require a different deployment strategy than one that provides a smaller improvement at much lower cost.

What to Build Next

For a new AI application, create the smallest working version first and establish a baseline with real evaluation examples. Improve the prompt before introducing infrastructure, add RAG when the application needs external or changing knowledge, and consider fine-tuning only after you can demonstrate a persistent behavior problem that simpler methods cannot solve. AWS's current customization guidance follows the same basic principle: increase customization as the requirements justify it rather than starting with the most complicated training approach.

The next practical step is to create a small evaluation dataset before changing your model. Once you can measure exactly where the baseline fails, the choice between prompting, retrieval, supervised fine-tuning, and reinforcement fine-tuning becomes an engineering decision based on evidence rather than guesswork.

S

Written by

Sarah Khan

I’m fascinated by artificial intelligence and the rapid changes happening around AI tools, models, and agents. I enjoy testing new AI technologies, following important developments, and understanding how they can be useful in real life. I like explaining complex AI topics in a simple and practical way.

51 posts published

All posts by this author

0 Comments

No comments yet. Be the first to share your thoughts.

Join the conversation

Log in or create a free account to leave a comment. You can edit or delete your own comments any time.