Lab: APIs and Web Interfaces
Due: Wednesday, December 2 at 11:59pm (one week after it is assigned) Worth: 8 points
Starter repository: github.com/rtealwitter/lab-fastapi
In this lab, you will put an LLM behind a web interface. Part 1 uses reddit’s API, an LLM API, and a small FastAPI server. Part 2 adds a web interface to a classmate’s LLM project. The final project also uses FastAPI.
Start
Select Use this template to create an independent lab-fastapi repository under your account, then clone that copy. Change into it and install its dependencies:
$ cd lab-fastapi
$ pip3 install -r requirements.txtPart 1: Web Servers and APIs
Using an API
An API (Application Programmer Interface) is just a webpage that returns JSON instead of HTML. The subreddit /r/ProgrammerHumor contains many of the memes shown in class, and its API returns the same content as JSON.
The endpoint is the URL that returns the JSON. For a reddit page you get the endpoint by adding .json to the end of the URL, so the endpoint for that subreddit is https://www.reddit.com/r/ProgrammerHumor.json. Open it in your browser and you will see a wall of JSON:
You do not need to know the meaning of every field. The API provides the same content reddit displays, without scraping.
For example, eBay provides a free API, and using it replaces all the scraping work from the earlier project with a single call to the API endpoint.
curl
curl is the standard shell tool for working with APIs, and all it does is download a webpage and print its contents to the screen. We can pull reddit’s JSON from the terminal with it:
$ curl https://www.reddit.com/r/ProgrammerHumor.jsonYou will see the same JSON printed in your terminal. The site https://cheat.sh/python is another handy example, a plain-text Python cheatsheet you can read in the browser or fetch with curl:
$ curl https://cheat.sh/pythonTalking to an LLM API
The Groq quickstart guide uses curl to call its LLM API:
curl -X POST "https://api.groq.com/openai/v1/chat/completions" \
-H "Authorization: Bearer $GROQ_API_KEY" \
-H "Content-Type: application/json" \
-d '{"model": "llama-3.3-70b-versatile", "messages": [{"role": "user", "content": "Explain the importance of fast language models"}]}'There is more inside this command than a URL. The -H headers handle authentication (logging in with your $GROQ_API_KEY) and declare that we are sending JSON, and the -d data carries the messages we are passing to the model. Copy the current command from the quickstart guide into your shell and run it, and you will get back a large JSON object with many fields; find the text the model responded with.
That raw object is hard to read, so pass it through Python’s built-in JSON formatter with a pipe:
$ curl <insert_curl_params_here> | python3 -m json.toolThe output starts something like this:
{
"id": "chatcmpl-c76ad728-221c-4d25-aa71-78272358c9b0",
"object": "chat.completion",
"created": 1777047642,
"model": "llama-3.3-70b-versatile",
"choices": [
{
"index": 0,
"message": {
"role": "assistant",
"content": "Fast language models are crucial for various applications..."This JSON has the same shape as the Python response object from your LLM project, where you read the reply with response.choices[0].message.content.
A Chat Web Interface
The URL you just used is https://api.groq.com/openai/v1/chat/completions. The openai in the path is there because Groq implements the OpenAI-compatible API. Many model providers and gateways support this API, allowing the same client tools to work with different models.
Custom chat interfaces can support direct conversation editing, configurable tools, and queries sent to several models at once. oobabooga and open-webui are two examples. We will use a simpler interface from gradio. The gradio_server.py file in the repo connects to any OpenAI-compatible endpoint, so we can point it straight at Groq:
$ python3 gradio_server.py --url=https://api.groq.com/openai/v1 --apikey=$GROQ_API_KEY
* Running on local URL: http://127.0.0.1:7860Visit the address it prints (probably http://127.0.0.1:7860), have a short conversation with the chatbot, and confirm everything works. The :7860 on the end of the address is a port: you will have several servers running at once, and the port says which one to connect to. A single computer has 2**16 = 65536 ports, so it can run that many servers at a time.
Your Own OpenAI-Compatible Endpoint
If we build our own OpenAI-compatible endpoint, the same tools can connect to our program.
The endpoint.py file is a small OpenAI-compatible endpoint written in FastAPI (the name comes from its being built for making APIs quickly). Run it:
$ python3 endpoint.py
INFO: Uvicorn running on http://127.0.0.1:8000 (Press CTRL+C to quit)Open the file and you will see it defines four routes, each a path you can connect to with curl or a browser: /, /spanish, /latin, and /v1/chat/completions. The first three just return a greeting:
$ curl http://127.0.0.1:8000/
hello world
$ curl http://127.0.0.1:8000/spanish
hola mundo
$ curl http://127.0.0.1:8000/latin
salve mundeThe fourth route is the chat endpoint, which you reach with a POST request carrying a message:
$ curl -X POST http://127.0.0.1:8000/v1/chat/completions -H "Content-Type: application/json" -d '{"messages":[{"role":"user","content":"hello"}]}'
{"id":"chatcmpl-123","object":"chat.completion","created":0,"model":"unknown","choices":[{"index":0,"message":{"role":"assistant","content":"this is response number 1"},"finish_reason":"stop"}],"usage":{"prompt_tokens":0,"completion_tokens":0,"total_tokens":0}}Put the same gradio interface in front of the new endpoint. Keep endpoint.py running in one terminal, then in a second terminal start gradio_server.py pointed at it:
$ python3 gradio_server.py --url=http://127.0.0.1:8000/v1
* Running on local URL: http://127.0.0.1:7860Visit the gradio address and you can chat with your own endpoint. This endpoint is a mock: it does not call an LLM at all and just returns a canned response no matter what you type. This tests the web connection, not the model.
Part 2: A Web Interface for a Classmate’s Project
Add a web interface to a classmate’s LLM project.
- Fork your partner’s project and clone it to your laptop. It does not matter which branch you use; the default branch has their working Project 3 code, and their Project 4 agent code lives on its assignment branch, so either gives you something to run.
- Copy
gradio_server.pyandendpoint.pyfrom this lab into the clone. - Modify
endpoint.pyto use your partner’sChatclass instead of the mock. The file is written to make this easy: you should only have to change theimportline so it imports theirChatrather than the one frommock_chat.py. - Test a conversation and confirm that a question like “what does the README say this project is about?” uses the project tools correctly. The grader exercises the endpoint directly, so no screenshot is needed.
- Commit and push your changes, then open a pull request to your partner’s project; your partner must accept it.
Submit
In your fork of your partner’s project, add a root-level submission.toml containing the accepted pull request URL:
pull_request_url = "https://github.com/partner/project/pull/123"No screenshot is required. The instructor may separately verify that the pull request was accepted.
Forking your partner’s repository in Part 2 is intentional: it lets you offer your change back through a pull request. Commit and sync the required files and submission.toml, then on Gradescope choose GitHub and submit your fork of your partner’s project and the branch containing your commit.