6.7 KiB
6.7 KiB
Ollama Vision Chat 🎥💬
A real-time camera-enabled chat application that connects to your local Ollama vision model, allowing you to have conversations with an AI that can see and understand what you're doing through your camera.
✨ Features
🎥 Real-time Camera Integration
- Continuous camera feed with live preview
- Automatic frame capture at configurable intervals (5-60 seconds)
- Manual photo capture with instant feedback
- Pause/resume auto-capture functionality
- Mobile and desktop camera support
💬 Interactive AI Chat
- Real-time messaging with vision-enabled AI
- Context-aware responses based on camera feed
- Message history with clear user/assistant distinction
- Auto-resizing text input for comfortable typing
- Smooth animations and modern UI
⚙️ Flexible Configuration
- Configurable Ollama server URL
- Support for multiple vision models (LLaVA, etc.)
- Adjustable capture intervals
- Real-time connection status monitoring
- Easy settings panel with instant updates
📱 Responsive Design
- Optimized for desktop and mobile devices
- Adaptive layout that works on any screen size
- Touch-friendly controls
- Modern glassmorphism design with smooth animations
🚀 Quick Start
Prerequisites
-
Install Ollama
# On macOS brew install ollama # On Linux curl -fsSL https://ollama.ai/install.sh | sh # On Windows # Download from https://ollama.ai/download -
Pull a Vision Model
# LLaVA (recommended) ollama pull llava # Or other vision models ollama pull llava:7b ollama pull llava:13b ollama pull bakllava -
Start Ollama Server
ollama serve
Installation & Usage
- Download the HTML file or clone this repository
- Open
index.htmlin a modern web browser - Grant camera permissions when prompted
- Wait for connection - the status indicator will turn green
- Start chatting! The AI can see what you're doing and respond accordingly
🛠️ Configuration
Settings Panel
Click the ⚙️ settings button to access:
- Ollama Server URL: Default
http://localhost:11434 - Vision Model: Default
llava(change to your preferred model) - Auto-capture Interval: 5-60 seconds between automatic photos
- Enable Auto-capture: Toggle automatic photo capturing
Supported Models
The app works with any Ollama vision model:
llava- General purpose vision modelllava:7b- Smaller, faster versionllava:13b- Larger, more capable versionbakllava- Alternative vision model- Custom vision models you've imported
💡 Usage Examples
What You Can Ask:
- "What am I holding in my hands?"
- "How does this setup look?"
- "Can you see what's on my screen?"
- "What color is my shirt?"
- "Help me organize this workspace"
- "What do you think of this drawing?"
- "Can you read this text for me?"
Perfect For:
- 🎨 Getting feedback on artwork or projects
- 📚 Reading assistance and text recognition
- 🏠 Home organization and decoration advice
- 👔 Fashion and styling feedback
- 🔧 Technical troubleshooting with visual context
- 🍳 Cooking assistance and recipe guidance
- 📖 Educational support with visual learning
🔧 Technical Details
Architecture
- Frontend: Pure HTML5, CSS3, and JavaScript (no frameworks)
- Camera: WebRTC getUserMedia API for camera access
- Image Processing: HTML5 Canvas for frame capture
- Communication: Fetch API for Ollama REST API calls
- UI: Modern CSS Grid and Flexbox layouts
Browser Requirements
- Modern browser with camera support
- JavaScript enabled
- HTTPS or localhost (required for camera access)
- WebRTC support
File Structure
ollama-vision-chat/
├── index.html # Main application file
├── README.md # This file
└── screenshots/ # (Optional) App screenshots
🚨 Troubleshooting
Common Issues & Solutions
🔴 "Failed to connect to Ollama"
- Ensure Ollama is running:
ollama serve - Check the server URL in settings (default:
http://localhost:11434) - Verify your firewall isn't blocking the connection
📷 "Failed to access camera"
- Grant camera permissions in your browser
- Ensure no other app is using the camera
- Try refreshing the page and granting permissions again
- Use HTTPS or localhost (camera requires secure context)
🤖 "Model not found"
- Install a vision model:
ollama pull llava - Check available models:
ollama list - Update the model name in settings to match your installed model
⏱️ "Slow responses"
- Try a smaller model like
llava:7b - Increase the auto-capture interval to reduce processing load
- Ensure your system has adequate RAM and CPU
📱 "Mobile camera issues"
- Grant camera permissions in browser
- Try both front and back cameras
- Ensure mobile browser supports WebRTC
🎯 Performance Tips
- Use smaller models (
llava:7b) for faster responses - Adjust capture intervals based on your use case
- Pause auto-capture when not needed to save resources
- Close other camera applications to avoid conflicts
- Use a stable internet connection for best performance
🔒 Privacy & Security
- All processing is local - your camera feed never leaves your device
- No data is stored - conversations are not saved
- Direct Ollama connection - no third-party services involved
- Open source - you can review all code in the HTML file
🤝 Contributing
Contributions are welcome! Here are some ideas:
- 🎨 UI/UX improvements
- 🔧 Additional model support
- 📱 Enhanced mobile experience
- 🌐 Multi-language support
- 📊 Usage analytics dashboard
- 🔄 Conversation export/import
📋 Roadmap
- Voice input/output support
- Multiple camera sources
- Image annotation tools
- Conversation export
- Custom model training integration
- Screen sharing capability
- Multi-user support
📜 License
This project is open source and available under the MIT License.
⭐ Support
If you find this project helpful:
- ⭐ Star this repository
- 🐛 Report issues and bugs
- 💡 Suggest new features
- 🤝 Contribute improvements
Built with ❤️ for the Ollama community
Have questions? Open an issue or start a discussion!