Skip to content

About

Local AI-powered Image Captioning app built using Hugging Face Transformers, BLIP, PyTorch, and Streamlit. Explores multimodal AI, local inference, transformer pipelines, image embeddings, autoregressive generation, and AI systems architecture using fully local vision-language models.

Topics

Resources

Stars

0 stars

Watchers

0 watching

Forks

Latest commit

 

History

10 Commits

Folders and files

NameName
Last commit message
Last commit date
 
 
 
 
 
 
 
 
 
 
 
 

Repository files navigation

Image Captioner (Local Vision AI App)

A local AI-powered image captioning application built using Hugging Face Transformers, BLIP, and Streamlit.

This project allows users to upload an image and generate intelligent captions completely locally without using cloud APIs like OpenAI or Gemini.

The goal of this project was to deeply understand:

  • Local AI inference
  • Vision-language models
  • Hugging Face ecosystem
  • Transformer pipelines
  • Multimodal AI systems
  • Model loading and caching
  • Inference latency and warm starts
  • AI system architecture

Features

  • Upload images through a Streamlit UI

  • Generate AI-powered image captions locally

  • Adjustable generation settings:

    • Temperature
    • Beam Search
    • Max Tokens
  • Inference latency measurement

  • Warm inference support

  • Fully local transformer inference

  • Modular backend architecture


Tech Stack

Frontend

  • Streamlit

AI / ML

  • Hugging Face Transformers
  • BLIP (Salesforce/blip-image-captioning-base)
  • PyTorch

Image Processing

  • Pillow (PIL)

Project Structure

ai-image-captioner/
│
├── app.py
│
├── utils/
│   └── caption_generator.py
│
├── sample_images/
│
└── requirements.txt

How It Works

Image Upload
↓
Image Preprocessing
↓
BLIP Vision Encoder
↓
Image Embeddings
↓
Transformer Decoder
↓
Caption Generation
↓
Display Result

The application uses a local BLIP transformer model to:

  • analyze uploaded images
  • extract visual embeddings
  • generate captions token-by-token using autoregressive decoding

Concepts Explored

This project was focused heavily on AI engineering concepts and system understanding, including:

  • Transformers
  • Attention Mechanism
  • Encoder-Decoder Architecture
  • Tokenization
  • Embeddings
  • Autoregressive Generation
  • Multimodal AI
  • Local Inference
  • Warm vs Cold Inference
  • Hugging Face Model Loading
  • AI Inference Pipelines
  • Beam Search
  • Generation Parameters
  • Model Caching
  • AI Infrastructure Tradeoffs

Installation

1. Clone the Repository

git clone https://github.com/decoded15/image-captioner.git
cd ai-image-captioner

2. Create Virtual Environment

python -m venv venv

Activate virtual environment:

Windows

venv\Scripts\activate

Mac/Linux

source venv/bin/activate

3. Install Dependencies

pip install -r requirements.txt

Run the App

streamlit run app.py

Model Used

BLIP

Model:

Salesforce/blip-image-captioning-base

BLIP (Bootstrapping Language-Image Pretraining) is a vision-language transformer model capable of generating image captions using multimodal transformer architectures.


Learning Outcomes

This project helped in understanding:

  • How local AI models run
  • How Hugging Face pipelines work
  • How multimodal transformers process images
  • How image embeddings are generated
  • How token generation works internally
  • How transformer decoding behaves
  • How AI inference systems are architected

Future Improvements

Potential upgrades for future versions:

  • Vision Question Answering (VQA)
  • OCR Integration
  • Object Detection
  • Webcam Support
  • Realtime Vision Analysis
  • Local Multimodal Chat Systems
  • Better Instruction-Tuned Vision Models
  • Quantized Model Support

Built by Dibyansh

About

Local AI-powered Image Captioning app built using Hugging Face Transformers, BLIP, PyTorch, and Streamlit. Explores multimodal AI, local inference, transformer pipelines, image embeddings, autoregressive generation, and AI systems architecture using fully local vision-language models.

Topics

Resources

Stars

0 stars

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages