Hadoop best performs on a cluster of multiple nodes/servers, however, it can run perfectly on a single machine, even a Mac, so we can use it for development. Also, Spark is a popular tool to process data in Hadoop. The purpose of this blog is to show you the steps to install Hadoop and Spark on a Mac.
Operating System: Mac OSX Yosemite 10.11.3
Hadoop Version 2.7.2
Spark 1.6.1
Hadoop Version 2.7.2
Spark 1.6.1
Pre-requisites
Summary Spark AR Studio Mac OS 10.12. This is a Beta product at the time of review. The producers provide a comprehensive guide for software training using the software product. The site provides tutorial, download software and full documents to get anyone started. From the site, the following example of work is provided. Install Latest Apache Spark on Mac OS. Following is a detailed step by step process to install latest Apache Spark on Mac OS. We shall first install the dependencies: Java and Scala. To install these programming languages and framework, we take help of Homebrew and xcode-select.
1. Install Java
Open a terminal window to check what Java version is installed.
$ java -version
$ java -version
- Download Spark for Mac - Desktop email client which enables you to categorize your messages into smart boxes, quickly clean up your inbox, and snooze messages to deal with them later.
- CleanSpark for Mac. 28 downloads Updated: September 30, 2020 Trial. Review Free Download specifications 100% CLEAN report malware. Quickly hide all desktop icons, folders, and active applications, and disable notifications for presentations, streaming, or any activity that requires.
If Java is not installed, go to https://java.com/en/download/ to download and install latest JDK. If Java is installed, use following command in a terminal window to find the java home path
$ /usr/libexec/java_home
$ /usr/libexec/java_home
Next we need to set JAVA_HOME environment on mac
$ echo export “JAVA_HOME=$(/usr/libexec/java_home)” >> ~/.bash_profile
$ source ~/.bash_profile
$ echo export “JAVA_HOME=$(/usr/libexec/java_home)” >> ~/.bash_profile
$ source ~/.bash_profile
2. Enable SSH as Hadoop requires it.
Go to System Preferences -> Sharing -> and check “Remote Login”.
Generate SSH Keys
$ ssh-keygen -t rsa -P “”
$ cat ~/.ssh/id_rsa.pub >> ~/.ssh/authorized_keys
$ ssh-keygen -t rsa -P “”
$ cat ~/.ssh/id_rsa.pub >> ~/.ssh/authorized_keys
Open a terminal window, and make sure we can do this.
$>ssh localhost
$>ssh localhost
Download Hadoop Distribution
Download the latest hadoop distribution (2.7.2 at the time of writing)
http://www.apache.org/dyn/closer.cgi/hadoop/common/hadoop-2.7.2/hadoop-2.7.2.tar.gz
http://www.apache.org/dyn/closer.cgi/hadoop/common/hadoop-2.7.2/hadoop-2.7.2.tar.gz
Create Hadoop Folder
Open a new terminal window, and go to the download folder, (let’s use “~/Downloads”), and find hadoop-2.7.2.tar
$ cd ~/Downloads
$ tar xzvf hadoop-2.7.2.tar
$ mv hadoop-2.7.2 /usr/local/hadoop
$ tar xzvf hadoop-2.7.2.tar
$ mv hadoop-2.7.2 /usr/local/hadoop
Hadoop Configuration Files
Go to the directory where your hadoop distribution is installed.
$ cd /usr/local/hadoop
$ cd /usr/local/hadoop
Then change the following files
$ vi etc/hadoop/hdfs-site.xml
2 4 6 | < property > < value >1</ value > </ configuration > |
$ vi etc/hadoop/core-site.xml
2 4 6 | < property > < value >hdfs://localhost:9000</ value > </ configuration > |
$ vi etc/hadoop/yarn-site.xml
2 4 6 | < property > < value >mapreduce_shuffle</ value > </ configuration > |
$ vi etc/hadoop/mapred-site.xml
2 4 6 | < property > < value >yarn</ value > </ configuration > |
Start Hadoop Services
Format HDFS
$ cd /usr/local/hadoop
$ bin/hdfs namenode -format
$ cd /usr/local/hadoop
$ bin/hdfs namenode -format
Start HDFS
$ sbin/start-dfs.sh
$ sbin/start-dfs.sh
Start YARN
$ sbin/start-yarn.sh
$ sbin/start-yarn.sh
Validation
Check HDFS file Directory
$ bin/hdfs dfs -ls /
$ bin/hdfs dfs -ls /
If you don’t like to include the bin/ every time you run a hadoop command, you can do the following
$ vi ~/.bash_profile
append this line to the end of the file “export PATH=$PATH:/usr/local/hadoop/bin”
$ source ~/.bash_profile
append this line to the end of the file “export PATH=$PATH:/usr/local/hadoop/bin”
$ source ~/.bash_profile
Now try to add the following two folders in HDFS that is needed for MapReduce job, but this time, don’t include the bin/.
$ hdfs dfs -mkdir /user
$ hdfs dfs -mkdir /user/{your username}
$ hdfs dfs -mkdir /user/{your username}
You can also open a browser and access Hadoop by using the following URL
http://localhost:50070/
http://localhost:50070/
Next: Spark
Installing Spark is a little easier. You can download the latest Spark here:
http://spark.apache.org/downloads.html
http://spark.apache.org/downloads.html
It’s a little tricky on choosing which package type. We want to choose “pre-build with user provided Hadoop [can use with most Hadoop distributions]” type, and the downloaded file name is spark-1.6.1-bin-without-hadoop.tgz
After spark is downloaded, we need to untar it. Open a terminal window and do the following:
$ cd ~/Downloads
$ tar xzvf spark-1.6.1-bin-without-hadoop.tgz
$ mv spark-1.6.1-bin-without-hadoop /usr/local/spark
$ tar xzvf spark-1.6.1-bin-without-hadoop.tgz
$ mv spark-1.6.1-bin-without-hadoop /usr/local/spark
Add spark bin folder to PATH
$ vi ~/.bash_profile
append this line to the end of the file “export PATH=$PATH:/usr/local/spark/bin”
$ source ~/.bash_profile
append this line to the end of the file “export PATH=$PATH:/usr/local/spark/bin”
$ source ~/.bash_profile
What about Scala?
Spark is written in Scala, so even though we can use Java to write Spark code, we want to install Scala as well.
Download Spark For Mac
Download Scala from here: http://www.scala-lang.org/download/
Choose the first one to download Scala in binary, and the downloaded file is scala-2.11.8.tar
Choose the first one to download Scala in binary, and the downloaded file is scala-2.11.8.tar
Untar Scala and move it to a dedicated folder
$ cd ~/Downloads
$ tar xzvf scala-2.11.8.tar
$ mv scala-2.11.8 /usr/local/scala
$ tar xzvf scala-2.11.8.tar
$ mv scala-2.11.8 /usr/local/scala
Download Apache Spark For Mac
Add Scala bin folder to PATH
$ vi ~/.bash_profile
append this line to the end of the file “export PATH=$PATH:/usr/local/scala/bin”
$ source ~/.bash_profile
append this line to the end of the file “export PATH=$PATH:/usr/local/scala/bin”
$ source ~/.bash_profile
Now you should be able to do the following to access Spark shell for Scala
$ spark-shell
Spark Email Windows
That’s it! Happy coding!