Function calling lets a language model move past text generation and act on the world: it parses a natural language request, works out the user's intent, and emits a structured payload naming the function to run along with the arguments it needs. Crucially, the model never executes anything itself. It selects the function, collects the required parameters, and returns them as JSON, which the host program deserializes and invokes in its own runtime.

To make this concrete, the worked example here is a Shopping Agent that helps users find fashion products and asks for clarification when the request is ambiguous. A prompt like "I'm looking for a shirt" or "Show me details about the blue running shirt" maps to either a keyword search API or a product-details lookup. All code samples are Python.

class ShoppingAgent:

    def run(self, user_message: str, conversation_history: List[dict]) -> str:
        if self.is_intent_malicious(user_message):
            return "Sorry! I cannot process this request."

        action = self.decide_next_action(user_message, conversation_history)
        return action.execute()

    def decide_next_action(self, user_message: str, conversation_history: List[dict]):
        pass

    def is_intent_malicious(self, message: str) -> bool:
        pass

The agent picks one of a predefined set of actions based on the user's message plus the conversation history, runs it, returns the result, and loops until the user's goal is met.

class Search():
    keywords: List[str]

    def execute(self) -> str:
        # use SearchClient to fetch search results based on keywords 
        pass

class GetProductDetails():
    product_id: str

    def execute(self) -> str:
 # use SearchClient to fetch details of a specific product based on product_id 
        pass

class Clarify():
    question: str

    def execute(self) -> str:
        pass

Building it test-first

Before implementing the full logic, unit tests pin down the expected behaviour:

def test_next_action_is_search():
    agent = ShoppingAgent()
    action = agent.decide_next_action("I am looking for a laptop.", [])
    assert isinstance(action, Search)
    assert 'laptop' in action.keywords

def test_next_action_is_product_details(search_results):
    agent = ShoppingAgent()
    conversation_history = [
        {"role": "assistant", "content": f"Found: Nike dry fit T Shirt (ID: p1)"}
    ]
    action = agent.decide_next_action("Can you tell me more about the shirt?", conversation_history)
    assert isinstance(action, GetProductDetails)
    assert action.product_id == "p1"

def test_next_action_is_clarify():
    agent = ShoppingAgent()
    action = agent.decide_next_action("Something something", [])
    assert isinstance(action, Clarify)

The decide_next_action function is built on OpenAI's API and a GPT model. It sends the user input and history to the model, then extracts the action type and any parameters:

def decide_next_action(self, user_message: str, conversation_history: List[dict]):
    response = self.client.chat.completions.create(
        model="gpt-4-turbo-preview",
        messages=[
            {"role": "system", "content": SYSTEM_PROMPT},
            *conversation_history,
            {"role": "user", "content": user_message}
        ],
        tools=[
            {"type": "function", "function": SEARCH_SCHEMA},
            {"type": "function", "function": PRODUCT_DETAILS_SCHEMA},
            {"type": "function", "function": CLARIFY_SCHEMA}
        ]
    )
    
    tool_call = response.choices[0].message.tool_calls[0]
    function_args = eval(tool_call.function.arguments)
    
    if tool_call.function.name == "search_products":
        return Search(**function_args)
    elif tool_call.function.name == "get_product_details":
        return GetProductDetails(**function_args)
    elif tool_call.function.name == "clarify_request":
        return Clarify(**function_args)

The call goes to OpenAI's chat completion API with a system prompt instructing gpt-4-turbo-preview to choose the action and pull out the parameters from the user's message and history. The structured JSON that comes back instantiates the matching action class, which in turn calls the underlying APIs such as search and get_product_details.

More details on the system prompt and function schemas are in part 2.

MCP and dynamic tool discovery

The Model Context Protocol (MCP), an open protocol proposed by Anthropic, is gaining traction as a standardized way to structure how LLM-based applications interact with the external world. A growing number of software-as-a-service providers are now exposing their services to LLM agents through it.

MCP defines a client-server architecture built from three components:

  • MCP Server — exposes data sources and tools (functions) that can be invoked over HTTP
  • MCP Client — manages communication between an application and the MCP Server
  • MCP Host — the LLM-based application (for example, a ShoppingAgent) that uses the data and tools from the MCP Server to accomplish a task; it reaches those capabilities through the MCP Client

The core problem MCP addresses is flexibility and dynamic tool discovery. In the shopping example, the set of available tools is hardcoded to three functions: search_products, get_product_details and clarify. That limits the agent's ability to adapt or scale to new types of requests, but it makes the agent easier to secure against malicious usage.

With MCP, the agent instead queries the MCP Server at runtime to discover which tools are available, then chooses and invokes the appropriate tool based on the user's query. This decouples the LLM application from a fixed set of tools and enables modularity, extensibility and dynamic capability expansion — valuable for complex or evolving agent systems. MCP adds complexity, but for some applications that complexity is justified. LLM-based IDEs or code generation tools, for instance, need to stay current with the latest APIs they can interact with. In theory, a general-purpose agent with access to a wide range of tools could handle a variety of user requests, unlike an agent confined to shopping tasks.

A simple MCP server for the shopping application looks like this. Note the GET /tools endpoint, which returns the list of all functions (tools) the server makes available.

TOOL_REGISTRY = {
    "search_products": SEARCH_SCHEMA,
    "get_product_details": PRODUCT_DETAILS_SCHEMA,
    "clarify": CLARIFY_SCHEMA
}

@app.route("/tools", methods=["GET"])
def get_tools():
    return jsonify(list(TOOL_REGISTRY.values()))

@app.route("/invoke/search_products", methods=["POST"])
def search_products():
    data = request.json
    keywords = data.get("keywords")
    search_results = SearchClient().search(keywords)
    return jsonify({"response": f"Here are the products I found: {', '.join(search_results)}"}) 

@app.route("/invoke/get_product_details", methods=["POST"])
def get_product_details():
    data = request.json
    product_id = data.get("product_id")
    product_details = SearchClient().get_product_details(product_id)
    return jsonify({"response": f"{product_details['name']}: price: ${product_details['price']} - {product_details['description']}"})

@app.route("/invoke/clarify", methods=["POST"])
def clarify():
    data = request.json
    question = data.get("question")
    return jsonify({"response": question})

if __name__ == "__main__":
    app.run(port=8000)

The corresponding MCP client handles communication between the MCP host (the ShoppingAgent) and the server:

class MCPClient:
    def __init__(self, base_url):
        self.base_url = base_url.rstrip("/")

    def get_tools(self):
        response = requests.get(f"{self.base_url}/tools")
        response.raise_for_status()
        return response.json()

    def invoke(self, tool_name, arguments):
        url = f"{self.base_url}/invoke/{tool_name}"
        response = requests.post(url, json=arguments)
        response.raise_for_status()
        return response.json()

Refactoring the ShoppingAgent as an MCP Host means it first retrieves the list of available tools from the MCP server and then invokes the appropriate function through the MCP client.

class ShoppingAgent:
    def __init__(self):
        self.client = OpenAI(api_key=os.getenv("OPENAI_API_KEY"))
        self.mcp_client = MCPClient(os.getenv("MCP_SERVER_URL"))
        self.tool_schemas = self.mcp_client.get_tools()

    def run(self, user_message: str, conversation_history: List[dict] = None) -> str:
        if self.is_intent_malicious(user_message):
            return "Sorry! I cannot process this request."

        try:
            tool_call = self.decide_next_action(user_message, conversation_history or [])
            result = self.mcp_client.invoke(tool_call["name"], tool_call["arguments"])
            return str(result["response"])

        except Exception as e:
            return f"Sorry, I encountered an error: {str(e)}"

    def decide_next_action(self, user_message: str, conversation_history: List[dict]):
        response = self.client.chat.completions.create(
            model="gpt-4-turbo-preview",
            messages=[
                {"role": "system", "content": SYSTEM_PROMPT},
                *conversation_history,
                {"role": "user", "content": user_message}
            ],
            tools=[{"type": "function", "function": tool} for tool in self.tool_schemas],
            tool_choice="auto"
        )
        tool_call = response.choices[0].message.tool_call
        return {
            "name": tool_call.function.name,
            "arguments": tool_call.function.arguments.model_dump()
        }
    
        def is_intent_malicious(self, message: str) -> bool:
            pass

Closing considerations

Function calling opens the door to novel user experiences and sophisticated agentic systems, but it also introduces new risks — especially when user input can ultimately trigger sensitive functions or APIs. Thoughtful guardrail design and proper safeguards can mitigate many of these risks. A prudent path is to enable function calling for low-risk operations first and extend it to more critical ones as the safety mechanisms mature.


Acknowledgements: thanks to Matteo Vaccari, Ben O'Mahony, Danilo Sato and Jim Gumbley for feedback on the early draft, and to Martin Fowler for providing the platform.

Significant revisions — 06 May 2025: initial publication.