Case Study
From 150 to 1,000 Items a Second: Real-Time GenAI Ingestion for Narrative Intelligence
How Peak Metrics, AWS, and AllCode replaced batch Kafka pipelines with real-time streaming, multi-modal embeddings, and RAG-ready vector search for disinformation defense.
Segment: Enterprise & Government | Use Case: Real-Time Narrative Intelligence, Multi-Modal GenAI Ingestion & RAG | Industry: Threat Intelligence & Media Monitoring

Real-Time Intelligence Challenge
Batch Pipelines Couldn't Keep Pace With Disinformation
Peak Metrics gives government agencies and enterprise clients an early-warning system against online disinformation, social media manipulation, and brand threat attacks, continuously processing unstructured, cross-channel media from news, social platforms, the web, and the dark web. Its legacy ingestion architecture relied on batch processing, Kafka message queues, and Dagster orchestration, which meant time-sensitive crisis signals showed up late.
Search was constrained to traditional keyword matching in Elasticsearch with no vector embeddings, which ruled out semantic search and automated narrative clustering. On the infrastructure side, Kubernetes ran through Porter and OpenSearch through Bonsai, third-party managed services that limited infrastructure control, pushed up EC2 compute costs, and created urgent migration pressure of their own.
High-profile government contracts were pushing in the other direction, demanding real-time streaming ingestion that could handle multi-modal assets, text, images, and long-form video, at meaningfully higher scale than the legacy pipeline was built for.
The Solution
A Real-Time, Multi-Modal GenAI Pipeline Built With AWS
Peak Metrics engaged AWS and AllCode to design and execute a 7-to-8-week proof of concept for a multi-modal data ingest, embeddings, and RAG enablement platform, replacing architectural debt with a pipeline built to unlock real-time narrative insight from day one.
Real-time streaming ingestion
Migrated ingestion off batch Kafka pipelines onto Amazon Kinesis Data Streams to decouple ingest from downstream processing and absorb high-volume burst traffic, with AWS Lambda consumer functions validating payload metadata and dispatching to workflow orchestrators.
Durable workflow orchestration
Temporal Cloud runs over AWS PrivateLink to meet SOC 2 and ISO compliance while keeping all traffic inside AWS network boundaries, with modality-specific pipelines for text, image, and video that guarantee exactly-once processing and automated retries.
Multi-modal embedding generation
Amazon Nova Multimodal Embeddings and Amazon Titan Text Embeddings v2 on AWS Bedrock generate vector representations for text, image, and video content, with embedding generation offloaded to managed Bedrock endpoints to keep GPU-intensive work out of the Kubernetes cluster.
Vector storage and RAG foundation
Self-managed Elasticsearch runs on a dedicated Amazon EKS cluster with compute (c6i) and memory (r6i) node groups, giving vector embeddings a shared foundation that powers RAG, semantic search, sentiment analysis, and proximity-based data visualization.
Operational hardening
Kinesis event-source mappings use bisect-on-error and capped retries so one bad message can't block a shard, client connections initialize at module scope to cut warm-invocation overhead, Bedrock calls fail over across regions, and Elasticsearch index lifecycle policies manage roughly 40GB of vector growth a day.
Technology
AWS-Native Stack Powering Real-Time Ingestion
At a target ingestion volume of 150 items per second (roughly 388.8 million documents a month), the team evaluated an all-Nova Multimodal configuration against a hybrid mix. All-Nova ran about $19,799 a month (~$237,595 ARR); pairing Titan Text v2 for text with Nova Multimodal for images and video cut that to about $12,605 a month (~$151,257 ARR). Peak Metrics also secured $200,000 in AWS funding through the Amazon Nova Multimodal Embeddings program to offset development and runtime costs.
The remediation and rollout stayed AWS-native throughout, engineering IaC delivered as Terraform with Argo CD automation for Kubernetes, and SOC 2/ISO compliance maintained via KMS encryption, TLS in transit, AWS PrivateLink connectivity, and IRSA role isolation.
Results
Before & After: From Batch Latency to Real-Time Intelligence
| Operational Metric | Legacy Infrastructure | Modernized GenAI Platform |
|---|---|---|
| Ingestion pipeline | Batch processing with Kafka queues and a Dagster orchestrator | Real-time streaming via Amazon Kinesis Data Streams and AWS Lambda |
| Search capabilities | Query-based text search with no vector embeddings | Multi-modal vector search and RAG across text, image, and video |
| Orchestration | Manual pipeline creation and batch orchestration | Durable, exactly-once execution via Temporal Cloud over PrivateLink |
| Embedding generation | No native embedding models in the ingestion path | External embedding generation via AWS Bedrock (Nova & Titan) |
| Infrastructure management | Third-party managed via Porter (EKS) and Bonsai (OpenSearch) | In-house, self-managed Elasticsearch on dedicated Amazon EKS clusters |
| Throughput target | Low-throughput batch mode, ~150 items/sec | Sustained real-time ingestion of up to 1,000 items/sec |
Business Impact
Measurable Impact: Real-Time GenAI at a Glance
The v0 embedding pipeline shipped to production with active vector indices already powering real-time narrative analysis, sustained sub-500ms P95 embedding latency, and end-to-end ingest-to-index times under 1,500ms for text, with a clear horizontal-scaling path to 1,000 items a second and a foundation in place for agentic workflows, smart category classification, and RAG-based chat inside customer workspaces.