gemini-vision
Implement Google Gemini API vision capabilities for image/document analysis including captioning, object detection, segmentation, and multi-image comparison.
Discover reusable agent skills, browse implementation details, and find the right skill for your workflow.
185 skills found
Implement Google Gemini API vision capabilities for image/document analysis including captioning, object detection, segmentation, and multi-image comparison.
Build stateful AI agents on Cloudflare Workers using the Agents SDK. Features real-time WebSockets, persistent state management, scheduled background tasks, and native tool integration for production-ready deployments.
Advanced TypeScript development agent: implements complex types, generics, branded types, and tRPC integration for end-to-end type safety.
Open-source infrastructure for reliable, multi-destination event delivery. Route webhooks to HTTP, SQS, RabbitMQ, Pub/Sub, EventBridge, or Kafka with built-in retries and observability.
Create professional technical diagrams and flowcharts using LaTeX TikZ with standardized Google Material and Anthropic-inspired design themes.
Expert assistant for designing and optimizing production-grade Trigger.dev background jobs, AI workflows, and resilient asynchronous task architectures in TypeScript.
Debug failing GitHub Actions CI checks by fetching logs, summarizing failures, and planning fixes.
Install and manage Codex agent skills from curated lists or GitHub repositories.
Build high-performance Solana apps with MagicBlock Ephemeral Rollups: sub-10ms latency, gasless transactions, and seamless integration for games and HFT.
Manage Tailwind CSS configuration, custom design tokens, themes, and design system utilities across the monorepo.
Design professional-grade brand identities using geometric primitives, negative space, and flat vector-style aesthetics via AI-driven branding logic.
Unified AI gateway for 100+ LLMs with OpenAI-compatible API, model fallbacks, load balancing, and enterprise-grade tools.