by Gerard Maas at DevoXX Belgium 2015
-
delivers a generic framework for parallel computing in a functional paradigm
-
fast growing ecosystem around
Spark Core
Spark SQL, dataframe API on top of it MLLib, ML out-of-the box GraphX, big data graph analytics Sparks Streaming, reuses spark to re-enable streaming application
Resilient Distributed Dataset unit of comutation in Spark a distributed collection as opposed to local collection like your daily programmatic Map or Lists they are immutable, memory-intensive and
Cachingis controllable
scalable fault-tolerant stream processing system
,---------------,
[ Kafka ]>| |>=Databases=>,--------------,
[ Flume ]>| [ Spark ] | | |
[ Kinesis ]>| [ Streaming ] |>=HDFS======>| |
[ Twitter ]>| | | Applications |
[ Sockets ]>| [ Spark ] |>=Server====>| |
[ HDFS/S3 ]>| | | |
[ Custom ]>| |>=Streams===>| |
'---------------' '--------------'
Micro batching is done over datastreams.
Every micro-batch will be an RDD.
On RDDs we can have Transformations
Then Actions will allow batches to get from Transformations to applications.
map, flatmap, filter
[% % % %] -> [& & & &]count, reduce, countByValue, reduceByKey[% % % %] -> nunion, join, cogroup[% %][$ $] -> [% % $ $]
-
Local
-
Standalone Cluster
-
Using a Cluster Manager (like mesos)
- within a
batch interval,receiversgatherblocksin parallel with ablock interval
bringing in the concept of partitions
- a formula for number of partitions,
receivers * batchInterval / blockInterval
rdd.cache() // cache RDD before iterating
Receiver-less model, Direct Kafka Stream
:) Simplified Psrallelism, Efficient, Exactly-once semantics :( Less degrees of freedom
Backpressure support
references:
virdata finding available via Wayback machine vanwilgenburg.wordpress.com/2015/02/15/spark-tuning-guide/ alvincjin.blogspot.in/2015/02/tuning-spark-streaming.html