Media & Entertainment / Social Automation AI-Powered Social Automation

Automating Movie Q&A on X with Hybrid RAG + Fine-Tuned LLMs

Built a production-ready X (Twitter) bot that answers movie questions from mentions using a modular AI layer that can switch between Local RAG, Hosted RAG, and a Fine-Tuned model. Delivered reliable ingestion, rate-limit resilience, and deployable endpoints for rapid iteration.

"Shipped a modular X bot that reliably replies to mentions with up-to-date movie answers—switching between Local RAG, Hosted RAG, and Fine-Tuned inference via config."
GenAIRAGFine-tuningFastAPIQdrantHuggingFaceOllamaX-APIPythonDocker
Backend
PythonFastAPIPostgreSQLQdrantDocker
AI Models
Mistral 7B InstructLlama 3.2 3B (fine-tuned)
Infrastructure
Hugging Face HubHugging Face Inference EndpointsGoogle Cloud (Vertex AI / Model Garden, Colab)Ollama (local inference)X API (Basic/Paid)
$4,500/mo avoided
Avoided real-time streaming API costs
Chose periodic mention polling + queueing instead of X Streaming API pricing.
15k
Monthly X API usage ceiling validated
Client reported dashboard usage (e.g., 168/15k), used to tune polling and rate-limit handling.
3 AI modes
Inference strategies available without rewrites
AI_ENV_TYPE routes between localRAG (Ollama), hosted RAG (HF/Qdrant), and fine-tuned model endpoints.
1 repo integration
Reduced handoff friction
Integrated working bot + AI modules into client’s FastAPI repo (finops360/twitter-llm-bot) with module README.

Problem Statement

The client needed a bot similar to interactive reply bots (e.g., @tomcruisebot-style) that could (1) ingest a CSV movie dataset, (2) answer questions from X mentions, and (3) stay current despite LLM knowledge cutoffs. Delivery was complicated by X API permission/rate-limit constraints, shifting deployment requirements (local → Hugging Face → Google Cloud), and a fast-moving scope that demanded production-grade modularization and environment separation.

Our Approach

Engineered a Python/FastAPI-based bot + AI services architecture with a clean module boundary: twitter_bot for X integration and ai_models for intelligence. Implemented a RAG pipeline backed by Qdrant (local via Docker and cloud) with ingestion utilities and an API surface (/ingest, tweet endpoints, health checks). Added configuration-driven AI routing (AI_ENV_TYPE) to switch between localRAG (Ollama + local Qdrant), hosted RAG, and fine-tuned inference endpoints. Hardened the bot with mention polling, queue management, and rate-limit handling to operate within X API constraints without requiring costly streaming APIs.

Hybrid AI Runtime (RAG + Fine-Tuned + Web Search toggle)

Technical Details
Implemented retrieval-augmented generation using Qdrant vector search over a CSV-derived knowledge base and local inference through Ollama (Mistral 7B) for a free/low-cost path. Built fine-tuning workflow using a Colab notebook and published the resulting model artifacts to Hugging Face Hub, then exposed them through Hugging Face Inference Endpoints. Optimized training cost by switching to a smaller base model (Llama 3.2 3B) when appropriate. Added a configuration switch (AI_ENV_TYPE) to route requests between local RAG, hosted RAG, and fine-tuned inference without code changes; centralized prompt templates and AI filtering under a dedicated ai_models module.
Business Value
Enabled the client to demo and iterate quickly with multiple AI strategies—local for development, hosted for production, and fine-tuned for consistent persona/format—while preserving a single bot codebase. This reduced deployment risk, improved maintainability for their team, and ensured the bot could answer both evergreen and newly-updated movie questions by choosing the best AI mode per environment and budget.

Challenges We Solved

X API permissions, rate limits, and non-streaming architecture

The bot needed to respond to mentions, but streaming access was unavailable/too expensive and the project hit unexpected rate limits even with low usage, blocking reliable testing.

Implemented periodic mention polling with robust rate-limit handling, queue management, and safe test-run cleanup. Validated plan/usage with the client, avoided streaming API dependency, and adjusted bot logic to remain functional under restrictive limits.

X APIPythonFastAPI

RAG ingestion reliability with change management

Client required vector DB reset/reinsert behavior and later asked for incremental updates from CSV changes, plus environment-driven tuning variables.

Built a dedicated ingestion script and ingestion API endpoint (/ingest) using Qdrant collections; added CSV hashing/state tracking and later moved toward centralized state tables. Parameterized key controls via env/config and documented the ingestion flow for production use.

QdrantDockerPythonFastAPI

Modular refactor into production FastAPI framework

Initial implementation needed to be merged into a pre-existing FastAPI framework with strict structure requirements (single app entrypoint, routers, separate AI modules, multi-env files).

Refactored codebase to move AI services out of twitter_bot into a standalone ai_models module, added APIRouter endpoints (tweeting, stats, ingestion, AI), relocated settings into app/config, and introduced .env.local/.env.dev/.env.prod patterns. Added README documentation per module to speed client KT and testing.

FastAPIAPIRouterPythonGitHub

Fine-tune deployment to Hugging Face + endpoint publishing

Client required a fine-tuned model that could be served via URL; access gating and account permissions caused repeated delays and confusion.

Delivered a working Colab fine-tuning notebook, published fine-tuned model versions to Hugging Face Hub (e.g., Rama_Movies4/Rama_Movies5), and configured Hugging Face Inference Endpoints for production-style API consumption. Guided the client through required 'Agree & access' steps for base models (Mistral/Llama) and provided the final push-to-hub block.

Google ColabHugging Face HubHugging Face Inference EndpointsTransformers

Project Timeline

1

Discovery

Aligned on a bot that monitors X mentions and replies with movie answers. Chose open-source LLM + RAG for cost control, clarified deployment expectations (Hugging Face + Google Cloud), and identified X API plan requirements for posting/replying.

2

Build

Delivered RAG API and ingestion pipeline (CSV → embeddings → Qdrant). Implemented mention monitoring, response posting, queue-based reliability, and rate-limit handling. Built fine-tuning workflow, published model artifacts and endpoints, and refactored into a modular FastAPI architecture with routers, env-based configuration, and documentation.

3

Launch

Integrated the bot into the client’s production repo, validated localRAG mode (Ollama + local Qdrant via Docker), confirmed hosted RAG/fine-tuned endpoints, and supported final testing with production X credentials (e.g., @askmoviebot). Delivered module READMEs and a high-level flow diagram for handover.

Want results like this for your business?

No pressure, no pitch deck. Just a free strategy call to explore what’s possible for your specific goals — and whether we’re the right fit to get you there.

Book a Free Strategy Call