Howdy,
Minutes from today's meeting follow below.
Had a really interesting conversation - though I think we didn't quite get through much of the agenda :) We did discuss some categories, things already packaged, etc. - I'm on duty to make a sub-wiki-page so we can collectively document some of these things.
Also had some talk around packaging Apache Bigtop as well as the intersections with that community - as many of their folks apparently use Fedora (YAY).
Anyway: Keep your intros, thoughts coming - I suspect we'll be gathering more humans in the coming weeks.
Thanks for coming!
-Robyn
Minutes: http://meetbot.fedoraproject.org/fedora-meeting-1/2013-03-07/big_data_sig.20... Full logs: http://meetbot.fedoraproject.org/fedora-meeting-1/2013-03-07/big_data_sig.20...
=============================== #fedora-meeting-1: Big Data SIG ===============================
Meeting started by rbergeron at 16:59:44 UTC. The full logs are available at http://meetbot.fedoraproject.org/fedora-meeting-1/2013-03-07/big_data_sig.20... .
Meeting summary --------------- * Who's around for fun? (rbergeron, 17:00:05) * present: rbergero, tflink (rbergeron, 17:01:15) * present: witlessb (rbergeron, 17:01:45) * present: threebean, zoglesby, samkottler, jsmith (rbergeron, 17:02:18)
* Agenda for today's first meeting :D (rbergeron, 17:03:49) * LINK: http://lists.fedoraproject.org/pipermail/bigdata/2013-March/000003.html (rbergeron, 17:04:32) * Agenda looks like: What this is all about, what do we have, what don't we have, what is anyone here interested in doing :) (rbergeron, 17:05:09)
* What's the Big Data SIG all about? (rbergeron, 17:05:57) * loosely quoting from o'reilly: "If the size of your data is part of the problem, it's Big Data." (rbergeron, 17:07:57) * IDEA: one part of it is getting a decent setup; other part is understanding tools and approaches needed (rbergeron, 17:11:42) * IDEA: two pieces you hear most about in Big Data seem to be massive storage, & parallel computing (hadoop, column databases, etc) (rbergeron, 17:12:16) * IDEA: another component is online processing or online analysis - predicting what is trending before its had time to hit disk (rbergeron, 17:13:29) * IDEA: for ex. - idea that google can predict flu outbreaks faster than public health agencies by watcihng search terms; financial tools as well apply to concept (rbergeron, 17:15:49) * IDEA: another ex. of stream processing is twitter analytics - looking for emerging topics in twitter streams (rbergeron, 17:16:52)
* What are the buckets or categories, and what do we have? (rbergeron, 17:18:21) * IDEA: orchestration, batch processing, stream processing are categories that come to mind - orch (zookeeper), batch (hadoop, disco), stream (storm) (rbergeron, 17:21:23) * IDEA: storage is another category (rbergeron, 17:22:25) * IDEA: full hadoop stack seems to be thought of as useful foundation layer for some types of work, but HDFS is getting attention as weak spot, with nosql dbs and gluster being used as alternatives (rbergeron, 17:23:00) * hdfs is part of the hadoop project; one can install hdfs without using the mapreduce part (rbergeron, 17:28:05) * lots of java libraries, servers like tomcat (used by solr, oozie) are already in (rbergeron, 17:44:01) * IDEA: we have pandas (not the animal, http://pandas.pydata.org) - useful for data analysis, not really big data (rbergeron, 17:49:40) * ACTION: rbergeron to add a sub-page of packges we have (unless someone beats me to it) (rbergeron, 17:50:18) * IDEA: interest in disco - seems to be more python-friendly than hadoop (though we are aware that there are python wrappers for hadoop) (rbergeron, 18:08:03) * spring is packaged - could package spring-hadoop (rbergeron, 18:11:24) * ACTION: bmahe to expound on openjdk/bug filing, as well as the wide world of bigtop packaging, as time permits :) (rbergeron, 18:12:51)
* Operation Agenda: Yeah... (rbergeron, 18:15:50) * ACTION: rbergeron to prod in meeting notes to get people to talk re: what would we like to do (we==they) (rbergeron, 18:23:54)
Meeting ended at 18:26:13 UTC.
Action Items ------------ * rbergeron to add a sub-page of packges we have (unless someone beats me to it) * bmahe to expound on openjdk/bug filing, as well as the wide world of bigtop packaging, as time permits :) * rbergeron to prod in meeting notes to get people to talk re: what would we like to do (we==they)
Action Items, by person ----------------------- * bmahe * bmahe to expound on openjdk/bug filing, as well as the wide world of bigtop packaging, as time permits :) * rbergeron * rbergeron to add a sub-page of packges we have (unless someone beats me to it) * rbergeron to prod in meeting notes to get people to talk re: what would we like to do (we==they) * **UNASSIGNED** * (none)
People Present (lines said) --------------------------- * rbergeron (163) * bmahe (59) * tflink (36) * ctyler (27) * threebean (8) * zodbot (5) * witlessb (4) * samkottler (1) * zoglesby (1) * jsmith (1)
Generated by `MeetBot`_ 0.1.4
.. _`MeetBot`: http://wiki.debian.org/MeetBot
On 03/07/2013 10:40 AM, Robyn Bergeron wrote:
Howdy,
Minutes from today's meeting follow below.
Had a really interesting conversation - though I think we didn't quite get through much of the agenda :) We did discuss some categories, things already packaged, etc. - I'm on duty to make a sub-wiki-page so we can collectively document some of these things.
Also had some talk around packaging Apache Bigtop as well as the intersections with that community - as many of their folks apparently use Fedora (YAY).
Anyway: Keep your intros, thoughts coming - I suspect we'll be gathering more humans in the coming weeks.
Thanks for coming!
-Robyn
Minutes: http://meetbot.fedoraproject.org/fedora-meeting-1/2013-03-07/big_data_sig.20... Full logs: http://meetbot.fedoraproject.org/fedora-meeting-1/2013-03-07/big_data_sig.20...
Hi,
I am Bruno and am very interested in cloud and big data technologies.
As promised (although a little bit later than planned), I am following up on the discussion from the irc meeting.
One of the project I work on intersects closely with this SIG. This project is Apache Bigtop (http://bigtop.apache.org/). The goal of Apache Bigtop is three folds: 1/ Provide top notch packages for Apache Hadoop related projects 2/ Provide a point of integration and testing for all these projects 3/ Provide means to reliably deploy a complete stack.
Apache Bigtop was donated by Cloudera to the Apache Foundation and is now the upstream of CDH (Cloudera's distribution), ubuntu hadoop packages (https://launchpad.net/~hadoop-ubuntu) and HDP 1.X from Hortonworks (haven't checked 2.X and not sure to which extend they have been modified).
We also use a lot of fedora into Apache Bigtop and I believe this SIG and Apache Bigtop could benefit from each others at least in some areas. For instance, Apache Bigtop provides live USB/CD images of Fedora with Apache Hadoop (and a bunch of other projects) pre-installed. We also use boxgrinder to build Centos VMs from a fedora build slave.
Right now, the list of projects Apache Bigtop supports is: * jsvc * tomcat * bigtop-utils -> Misc. tools, such asauto-detecting a JVM * crunch -> Library for writing/testing/running MapReduce pipelines * datafu -> Collection of libraries for pig * flume -> distributed, reliable, and available service for efficiently collecting, aggregating, and moving large amounts of log data * giraph -> Graph processing on top of Apache Hadoop * hadoop * hama -> Apache Hama is a pure BSP (Bulk Synchronous Parallel) computing framework * hbase -> columnar database * hive -> Hive is a data warehouse system for Hadoop that facilitates easy data summarization, ad-hoc queries, and the analysis of large datasets stored in Hadoop compatible file systems. Enables to run SQL-like queries * hue -> browser-based desktop interface for interacting with Apache Hadoop * mahout -> Library for machine learning and data mining on top of Apache Hadoop * oozie -> workflow scheduling * pig -> High level language to process data * solr -> search platform * sqoop -> Tool to transfert data between Apache Hadoop and relational databases * whirr -> set of libraries for running cloud services * zookeeper -> server which enables highly reliable distributed coordination * HCatalog -> In process of integration into Apache Bigtop. This is a service to manage data's metadata
All these projects have packages for Debian/Ubuntu/SLES11/Fedora/CentOS. They also have tests to exercise integration points between all of them (ex: Hive can use HBase which sits on top of HDFS). And in order to run these tests, we also have a test framework. Also before we can test for integration, we also have to ensure they can be properly installed/upgraded/removed, with the right users, ulimits, rights and so forth. So to that end, we also have a large chunk of the tests and testing framework dedicated to testing the packages themselves.
And finally, regarding the deployment, we have the following: * Boxgrinder recipe people can use and modify to suit their need * kickstart file to build a live fedora image * puppet recipes to deploy all these services. I am not sure if committed it, but I also had some puppet recipes which would install/setup and integrate all these services with ganglia and nagios automatically. We routinely use these recipes to automatically deploy a cluster on ec2, run tests and get tests reports
From my experience with these projects, some of the pain points fedora would have in integrating these packages are: * None of them check for openjdk compatibility. They would welcome patches, they would love to support openjdk, but no one has had the time or resources to ensure compatibility. Same apply for (open)JDK 7 * All these projects move fast and can have quite a bit of dependencies (which can also change over time). Note also that most of them use maven (so at least it would be uniform). In Apache Bigtop we side stepped this issue by not packaging dependencies separately. We had to make a choice between supporting more distributions or packaging the dependencies and we picked the former. * All these projects just ask users to disable selinux. So there is no integration with selinux at this time * Apache Hadoop also had issues with ipv6. So they used to ask users to disable it. I am not sure if this is still true.
I hope this gives some overview, but feel free to ask any question.
Thanks, Bruno
On 03/09/2013 05:35 PM, Bruno Mahé wrote:
On 03/07/2013 10:40 AM, Robyn Bergeron wrote:
Howdy,
Minutes from today's meeting follow below.
Had a really interesting conversation - though I think we didn't quite get through much of the agenda :) We did discuss some categories, things already packaged, etc. - I'm on duty to make a sub-wiki-page so we can collectively document some of these things.
Also had some talk around packaging Apache Bigtop as well as the intersections with that community - as many of their folks apparently use Fedora (YAY).
Anyway: Keep your intros, thoughts coming - I suspect we'll be gathering more humans in the coming weeks.
Thanks for coming!
-Robyn
Minutes: http://meetbot.fedoraproject.org/fedora-meeting-1/2013-03-07/big_data_sig.20...
Full logs: http://meetbot.fedoraproject.org/fedora-meeting-1/2013-03-07/big_data_sig.20...
Hi,
I am Bruno and am very interested in cloud and big data technologies.
As promised (although a little bit later than planned), I am following up on the discussion from the irc meeting.
One of the project I work on intersects closely with this SIG. This project is Apache Bigtop (http://bigtop.apache.org/). The goal of Apache Bigtop is three folds: 1/ Provide top notch packages for Apache Hadoop related projects 2/ Provide a point of integration and testing for all these projects 3/ Provide means to reliably deploy a complete stack.
Apache Bigtop was donated by Cloudera to the Apache Foundation and is now the upstream of CDH (Cloudera's distribution), ubuntu hadoop packages (https://launchpad.net/~hadoop-ubuntu) and HDP 1.X from Hortonworks (haven't checked 2.X and not sure to which extend they have been modified).
We also use a lot of fedora into Apache Bigtop and I believe this SIG and Apache Bigtop could benefit from each others at least in some areas. For instance, Apache Bigtop provides live USB/CD images of Fedora with Apache Hadoop (and a bunch of other projects) pre-installed. We also use boxgrinder to build Centos VMs from a fedora build slave.
Right now, the list of projects Apache Bigtop supports is:
- jsvc
- tomcat
- bigtop-utils -> Misc. tools, such asauto-detecting a JVM
- crunch -> Library for writing/testing/running MapReduce pipelines
- datafu -> Collection of libraries for pig
- flume -> distributed, reliable, and available service for efficiently
collecting, aggregating, and moving large amounts of log data
- giraph -> Graph processing on top of Apache Hadoop
- hadoop
- hama -> Apache Hama is a pure BSP (Bulk Synchronous Parallel)
computing framework
- hbase -> columnar database
- hive -> Hive is a data warehouse system for Hadoop that facilitates
easy data summarization, ad-hoc queries, and the analysis of large datasets stored in Hadoop compatible file systems. Enables to run SQL-like queries
- hue -> browser-based desktop interface for interacting with Apache Hadoop
- mahout -> Library for machine learning and data mining on top of
Apache Hadoop
- oozie -> workflow scheduling
- pig -> High level language to process data
- solr -> search platform
- sqoop -> Tool to transfert data between Apache Hadoop and relational
databases
- whirr -> set of libraries for running cloud services
- zookeeper -> server which enables highly reliable distributed
coordination
- HCatalog -> In process of integration into Apache Bigtop. This is a
service to manage data's metadata
All these projects have packages for Debian/Ubuntu/SLES11/Fedora/CentOS. They also have tests to exercise integration points between all of them (ex: Hive can use HBase which sits on top of HDFS). And in order to run these tests, we also have a test framework. Also before we can test for integration, we also have to ensure they can be properly installed/upgraded/removed, with the right users, ulimits, rights and so forth. So to that end, we also have a large chunk of the tests and testing framework dedicated to testing the packages themselves.
And finally, regarding the deployment, we have the following:
- Boxgrinder recipe people can use and modify to suit their need
- kickstart file to build a live fedora image
- puppet recipes to deploy all these services. I am not sure if
committed it, but I also had some puppet recipes which would install/setup and integrate all these services with ganglia and nagios automatically. We routinely use these recipes to automatically deploy a cluster on ec2, run tests and get tests reports
From my experience with these projects, some of the pain points fedora would have in integrating these packages are:
- None of them check for openjdk compatibility. They would welcome
patches, they would love to support openjdk, but no one has had the time or resources to ensure compatibility. Same apply for (open)JDK 7
- All these projects move fast and can have quite a bit of dependencies
(which can also change over time). Note also that most of them use maven (so at least it would be uniform). In Apache Bigtop we side stepped this issue by not packaging dependencies separately. We had to make a choice between supporting more distributions or packaging the dependencies and we picked the former.
- All these projects just ask users to disable selinux. So there is no
integration with selinux at this time
- Apache Hadoop also had issues with ipv6. So they used to ask users to
disable it. I am not sure if this is still true.
I hope this gives some overview, but feel free to ask any question.
Thanks, Bruno
Some other interesting projects are: * Apache S4 (incubating): Stream processing * Apache Helix (incubating): Cluster management framework * Apache bookkeeper (part of zookeeper) * Apache Kafka (incubating): scalable pubsub * Apache Blur (incubating): scalable search engine * Apache Cassandra * Apache nutch (the search engine which started the need for Apache Hadoop) * Apache Accumulo: data store based on the BigTable design
Aside from that, there are a bunch of various python wrappers for Apache Hadoop HDFS and/or mapreduce. There are also a lot of "glue" projects to tie all these projects together. For instance Input/Output modules between mongodb and Apache Hadoop.
In the context of this SIG, we could also take some use cases and dog food them with our own data. This would not only show how to put everything together but also give a great example of how Fedora can help you get your job done. Even if fedmsg does not generate petabytes of data, it could be interesting to document use cases of how to ingest that data and mine it. We may even end up with some interesting results.
Thanks, Bruno
bigdata@lists.fedoraproject.org