Sign inSign up

kaskadaio/jupyter

By kaskadaio

•Updated about 3 years ago

Kaskada pre-installed in a Jupyter notebook environment

Image
0

883

kaskadaio/jupyter repository overview

This repo provides an experimentation playground with Kaskada pre-installed in a Jupyter notebook environment. If you don't require the Jupyter environment consider using the smaller kaskadaio/engine⁠ container image instead.

⁠Kaskada: Modern, open-source event-processing

kaskada.io⁠ | Docs⁠

Kaskada is a unified event processing engine that provides all the power of stateful stream processing in a high-level, declarative query language designed specifically for reasoning about events in bulk and in real time.

Kaskada's query language builds on the best features of SQL to provide a more expressive way to compute over events. Queries are simple and declarative. Unlike SQL, they are also concise, composable, and designed for processing events. By focusing on the event-processing use case, Kaskada's query language makes it easier to reason about when things happen, state at specific points in time, and how results change over time.

Kaskada is implemented as a modern compute engine designed for processing events in bulk or real-time. Written in Rust and built on Apache Arrow, Kaskada can compute most workloads without the complexity and overhead of distributed execution.

Read more at kaskada.io⁠.

⁠Features

  • Stateful aggregations: Aggregate events to produce a continuous timeline whose value can be observed at arbitrary points in time.
  • Automatic joins: Every expression is associated with an “entity”, allowing tables and expressions to be automatically joined. Entities eliminate redundant boilerplate code.
  • Event-based windowing: Collect events as you move through time, and aggregate them with respect to other events. Ordered aggregation makes it easy to describe temporal interactions.
  • Pipelined operations: Pipe syntax allows multiple operations to be chained together. Write your operations in the same order you think about them. It's timelines all the way down, making it easy to aggregate the results of aggregations.
  • Row generators: Pivot from events to time-series. Unlike grouped aggregates, generators produce rows even when there's no input, allowing you to react when something doesn't happen.
  • Continuous expressions: Observe the value of aggregations at arbitrary points in time. Timelines are either “discrete” (instantaneous values or events) or “continuous” (values produced by a stateful aggregations). Continuous timelines let you combine aggregates computed from different event sources.
  • Native time travel: Shift values forward (but not backward) in time, allowing you to combine different temporal contexts without the risk of temporal leakage. Shifted values make it easy to compare a value “now” to a value from the past.
  • Simple, composable syntax: It is functions all the way down. No global state, no dependencies to manage, and no spooky action at a distance. Quickly understand what a query is doing, and painlessly refactor to make it DRY.

⁠Setting up your experimentation environment

To get started, you start a docker container that comes pre-installed with Jupyter and Kaskada:

docker run --rm -p 8888:8888 kaskadaio/jupyter

Then open the url specified in the logs from the docker run command in your browser.

Next, in the Jupyter environment, start a new python3 notebook, and then run the following code:

from kaskada.api.session import LocalBuilder

session = LocalBuilder().build()

%load_ext fenlmagic

Then continue with the Getting Started⁠ docs from Loading Data into a Table⁠

⁠Join Us!

We're building an active, inclusive community of users and contributors. Come get to know us on Slack⁠ - we'd love to meed you!

For specific problems, file an issue⁠.

Tag summary

Content type

Image

Digest

sha256:13ba396ed…

Size

1.3 GB

Last updated

about 3 years ago

docker pull kaskadaio/jupyter