This project is a LangGraph-powered chatbot agent deployed on Google Cloud Run. It demonstrates a clean agent architecture end-to-end: Streamlit for the UI, FastAPI for request/response, and LangGraph as the orchestrator that conditionally routes between tool execution and LLM generation.
🌐 Live demo: https://cloudrun-langraph-agent-499606806948.me-central1.run.app
At a high level:
- Streamlit collects the user’s message + model selection (UI only).
- FastAPI exposes
/chatand calls the agent (API only). - LangGraph routes the request through nodes (orchestration).
- OpenAI SDK is used by the LLM node (GPT-4o / 4o-mini).
The agent is a LangGraph state machine with three core nodes:
- router: inspects the latest message and sets
state["route"] - calc: runs a safe calculator tool when the message looks like math
- respond: calls the LLM and returns a natural-language reply
✅ Conditional routing happens in LangGraph: the router sets a route label, and LangGraph uses add_conditional_edges(...) to choose which node runs next.
+-------+ +--------+ +--------+
| Input | --> | router | --> | calc | --> END
+-------+ +--------+ +--------+
|
+---------> +---------+ --> END
| respond |
+---------+
def router(state: AgentState) -> AgentState:
last = state["messages"][-1]
text = (last.content or "").strip()
looks_math = any(ch.isdigit() for ch in text) and any(sym in text for sym in "+-*/%()^")
state["route"] = "calc" if looks_math else "respond"
return state
def build_graph():
g = StateGraph(AgentState)
g.add_node("router", router)
g.add_node("calc", do_calc)
g.add_node("respond", respond)
g.set_entry_point("router")
g.add_conditional_edges("router", _route, {"calc": "calc", "respond": "respond"})
g.add_edge("calc", END)
g.add_edge("respond", END)
return g.compile()LLM responses are generated with the OpenAI Python SDK using Chat Completions. The function supports a model parameter (defaulting to gpt-4o-mini) and keeps responses fast and controlled using a smaller max_tokens and low temperature.
def generate(prompt: str, model: str = "gpt-4o-mini") -> str:
key = os.getenv("OPENAI_API_KEY")
if not key:
raise RuntimeError("Missing OPENAI_API_KEY")
client = OpenAI(api_key=key)
resp = client.chat.completions.create(
model=model,
messages=[{"role": "user", "content": prompt}],
max_tokens=150,
temperature=0.2,
)
return resp.choices[0].message.content or ""Streamlit sends JSON to FastAPI /chat with the user’s message and selected model. FastAPI validates the request, calls invoke(...), and returns the final reply back to the UI.
def call_agent(text: str, model_name: str) -> str:
r = requests.post(
f"{API_BASE}/chat",
json={"message": text, "model": model_name},
timeout=60,
)
r.raise_for_status()
return r.json()["reply"]
@app.post("/chat", response_model=ChatResponse)
def chat(req: ChatRequest):
msg = (req.message or "").strip()
if not msg:
raise HTTPException(status_code=400, detail="message cannot be empty")
try:
out = invoke(msg, model=req.model)
reply = out["messages"][-1].content
return ChatResponse(reply=reply)
except Exception as e:
raise HTTPException(status_code=500, detail=str(e))(Local Runtime)
┌───────────────────────────────┐
│ Streamlit UI (streamlit_app.py)│
│ - message + model dropdown │
└───────────────┬───────────────┘
│ HTTP POST /chat
v
┌───────────────────────────────┐
│ FastAPI (app/main.py) │
│ - validates payload │
│ - calls LangGraph invoke(...) │
└───────────────┬───────────────┘
│ graph.invoke(state)
v
┌───────────────────────────────┐
│ LangGraph (app/agent/graph.py) │
│ router → calc OR respond │
└───────────────┬───────────────┘
│ (respond node)
v
┌───────────────────────────────┐
│ OpenAI Chat Completions (SDK) │
│ gpt-4o / gpt-4o-mini │
└───────────────────────────────┘
(GCP Deployment)
GitHub Repo
→ Cloud Build Trigger (build + push + deploy)
→ Cloud Run Service (public URL)
This project runs as a single Docker container, which makes local development and cloud deployment consistent.
- Dockerfile defines the runtime environment (Python + dependencies + start command).
- Cloud Build is triggered on pushes (CI/CD), and runs:
- Build the container image
- Push the image to a registry
- Deploy the latest revision to Cloud Run
- Cloud Run serves the container on a public HTTPS URL.
- Wire the model selection end-to-end (UI → FastAPI → LangGraph state → LLM node).
- Improve routing logic beyond heuristics (intent-based router).
- Add more tools/nodes (search, parsing, structured outputs).
- Add memory/thread persistence for multi-turn conversations.
- Improve UI error handling and response formatting.
- Python
- LangGraph (agent orchestration + conditional routing)
- LangChain Core (message objects)
- FastAPI (backend API)
- Uvicorn (ASGI server)
- Streamlit (UI)
- Requests (UI → API calls)
- OpenAI Python SDK (Chat Completions: GPT-4o / GPT-4o-mini)
- Docker (containerization)
- Google Cloud Build (CI/CD)
- Google Cloud Run (hosting)

