Course Content
Cribl Stream
1. Cribl Stream Architecture
- Overview and use cases
- Leader/Worker node model
- Deployment topologies:
- On-premises
- Hybrid
- Cloud
- Data flow concepts:
- Sources → Pipelines → Destinations
2. Cribl Setup & Workers Installation (Cloud-based)
- Cribl.Cloud onboarding
- Creating Worker Groups
- Deploying and registering Workers
- Sizing and scaling considerations
- Cloud vs. on-premises trade-offs
3. Cribl Admin Activities
- User management and RBAC
- Git-based configuration management
- Licensing and capacity management
- Environment promotion:
- Dev
- Stage
- Production
4. Sources and Destinations
- Configuring inbound sources:
- Splunk
- Syslog
- HTTP/S
- Kafka
- Files
- Configuring outbound destinations:
- S3
- Splunk
- Elastic
- Datadog
- Persistent queues
5. Cribl Integration
- Integration with:
- Splunk
- Datadog
- Elasticsearch
- ServiceNow
- BigPanda
- Authentication and connection setup
- End-to-end data flow validation
6. Routes and Pipelines
- Route design and precedence
- Conditional routing and filters
- Building and chaining pipelines
- Cloning and testing pipelines
7. Data Filtering and Transformation
- Core Functions:
- Eval
- Mask
- Parser
- Drop
- Aggregation
- Sampling
- Regex-based filtering
- Enrichment using lookups
8. Edge Agents
- Cribl Edge architecture
- Agent deployment and fleet management
- Source collection at the edge
- Edge vs. Stream Worker use cases
9. Search and Lake Fundamentals
- Cribl Search architecture
- Federated search across data-in-place
- Cribl Lake storage
- Retention and lifecycle policies
10. OpenTelemetry Integration
- OTel Collector basics
- Ingesting OTel:
- Traces
- Metrics
- Logs
- Building OTel-aware pipelines
- Exporting to observability backends
11. BigPanda Integration
- Configuring BigPanda as an alert/event destination
- Event correlation setup
- End-to-end validation and testing
12. Troubleshooting and Monitoring
- Cribl health monitoring dashboards
- Internal logs and notifications
- Common failure scenarios and resolution
- Capacity and performance planning
ClickHouse
1. ClickHouse Fundamentals
- Columnar database concepts
- ClickHouse architecture
- Key observability use cases:
- Logs
- Metrics
- Traces at scale
2. Installation and Deployment
- Single-node installation
- Cluster installation
- Docker-based setup
- Cloud vs. self-managed deployment
- Core configuration files:
config.xmlusers.xml
3. ClickHouse Admin Activities
- User management
- Roles and permissions
- Backup and restore
- Version upgrades
- Quotas and resource management
4. Data Types & Schema Design
- Native data types
- Schema design best practices
- Choosing primary keys
- Partitioning strategies
5. Table Engines
- MergeTree family
- ReplacingMergeTree
- SummingMergeTree
- AggregatingMergeTree
- Distributed table engine
6. Data Ingestion
- Batch vs. streaming ingestion
- Kafka table engine
- Native/JDBC/ODBC drivers
- Bulk insert best practices
7. Querying Data
- ClickHouse SQL fundamentals
- Joins and aggregations
- Filtering large datasets
- Sorting large datasets
8. Advanced SQL
- Window functions
- Materialized views
- Array and nested data functions
- Approximate aggregate functions
9. Performance Optimization
- Indexing strategies:
- Primary indexes
- Skip indexes
- Query profiling
EXPLAIN- Compression codecs
- Query tuning
10. Distributed ClickHouse
- Sharding and replication
- Cluster topology design
- ZooKeeper/Keeper coordination
- Failover scenarios
11. Monitoring & Administration
- System tables:
system.query_logsystem.metrics
- Grafana integration
- Alerting on cluster health
12. ClickHouse for Observability
- Using ClickHouse as a backend for:
- Logs
- Metrics
- Traces
- Comparison with other observability stores
- Schema patterns for telemetry data
13. Integrations – OpenTelemetry
- OTel exporter to ClickHouse
- Schema mapping for OTel data
- End-to-end pipeline demonstration
