# DatabeanStalk documentation

Welcome to DatabeanStalk official documentation. DatabeanStalk is a fully collaborative notebook where you can run python, PySpark, Scala and R in your browser.

The DatabeanStalk data science and machine learning plateform support fully managed Apache Spark which enables data teams to work together and solve world's toughest data problems.

![](/files/S1JollceKDjwLb5wGPCx)

DatabeanStalk's managed spark on kubernetes where you can initialize PySpark, Spark and SparkSQL context and DatabeanStalk manage master and workers executors for you.

DatabeanStalk support data scientist, data engineer and data analyst to work on following use cases.

* Run advance data analytics and machine learning complex problem at scale
* Make data warehousing easier using python data-frame and spark sql.
* Real-time detect threat analysis using spark streaming, data science libaries and Apache Airflow integration.
* Analyze time-series data in realtime or update dashboard chart.
* Scheduled based data extract translate and load using Apache Airflow integration

DatabeanStalk offer managed Apache Airflow, so user able to build, package and deploy machine learning jobs and also integrate any third party library with airflow.

Using Grafana dashboard, user able to monitor server resources and analyze application logs.

The following screenshot shows the DatabeanStalk landing page which has JupyterHub, Apache Airflow and Grafana. Using these technology many users able to solve world's complex data problems.


# Registration

DatabeanStalk offer free access to community users and 30 days free trial for enterprise user. DatabeanStalk community plateform is shared environment but it has absolutly private isolated workspace. DatabeanStalk business plateform is package where user able to deploy its own AWS account and its complete private workspace.

DatabeanStalk invite educational users and students to create free community based account and busness users to free enterprise 30 days trial account for Data Science and Machine Learning for deploy and train data science and prediction model using Python and Spark based online notebook.


# Onboard to DatabeanStalk

To onboard to DatabeanStalk, please follow the steps in this guide. In the following sections, you learn how to create your DatabeanStalk workspace account and sign in.

### <mark style="color:blue;">**Create a DatabeanStalk account**</mark>

Completes the following steps to create your DatabeanStalk workspce account.

1. Go to the [DatabeanStalk website](https://databeanstalk.com) and click on signup.
2. To get started with DatabeanStalk, you can choose between DatabeanStalk Free Trial and Free Community Edition.

![](/files/HHRRcX8zjdoDUrfAdYbc)

### <mark style="color:blue;">**DatabeanStalk Free Trial vs Free Community Edition**</mark>

DatabeanStalk Free Trial gives you full product access and it has more features and flexibility than Community edition. In this trial version, DatabeanStalk request your AWS account role base access and s3 bucket access for data storage and connectivity. This product will take up to 30 minutes to deploy in your AWS account and you will receive your own url link in registered email for product access.

Free Community Edition is complete free for user with limited resource access. User able to use  up to 2 CPU core and 6GB RAM and 15GB storage.


# Set up your DatabeanStalk account

Once you signed up your DatabeanStalk Free Trial, you call follow this guide to set up admin account and deployment.

1. For set up this process you need to create AWS account or you can use existing AWS account.
2. Once you set up your password from verification email you can [login](https://apps.databeanstalk.com/Identity/Account/Login?ReturnUrl=%2F) here using your email and newly created password.
3. Select the DatabeanStalk plan based on your business needs.

![](/files/4LVNWKSKh3HEDf2XlNJ3)

4\. Enter billing detail, DatabeanStalk will not charge till your free trial expire. Once your free trial completes, DatabeanStalk only charge subset of your AWS infrastructure billing.

![](/files/jhxNL45TYQXGaqeqRTRf)

5\. You can copy external ID and use it in your AWS account to create cross account role. you can follow this guide to create databeanstalk cross account role in your account.

![](/files/YYyua9zieSCas1UMSfvD)

6\. AWS storage: create your unque s3 bucket and use that bucket name here and click on Create Policy and databeanStalk will create a permission for you and you can copy and paste in your bucket resource policy.

![](/files/7Q9NeAtar24Xzlgxmgic)

7\. Deploy DatabeanStalk: This process take up to 30 minutes to deploy in your AWS account and you will get notification in your email address. Click on Deploy DatabeanStalk.

![](/files/YKXbNFbqLeAToYshO0Hp)


# DatabeanStalk Free Trial

DatabeanStalk Free Trial gives you full product access and it has more features and flexibility than Community edition. In this trial version, DatabeanStalk request your AWS account role base access and s3 bucket access for data storage and connectivity. This product will take up to 30 minutes to deploy in your AWS account and you will receive your own url link in registered email for product access.


# Sign up

### <mark style="color:blue;">**Sign up for DatabeanStalk Free Trial**</mark>

1. Click on Sign up and select DatabeanStalk Free Trial.
2. Enter you name, email, company, contact number and click on Get Started.
3. It will send verification email to your email accress.
4. You can open email inbox and click on verification link and set up password.

![](/files/RrIDeHCLjWs78F6nzNva)


# Login

<mark style="color:blue;">**Login into DatabeanStalk Free Trial**</mark>

1. [Login](https://apps.databeanstalk.com/Identity/Account/Login?ReturnUrl=%2F) to your DatabeanStalk Free Trial version.
2. Follow this [guide](/databeanstalk-guides/registration/set-up-your-databeanstalk-account) to set up your DatabeanStalk plan in trial version.
3. For the community user click on Get Started in free community edition and for business user click on Get Started in DatabeanStalk free trial version.
4. From the signup form, enter your firstname, lastname, email and company details.
5. Select on Get Started For Free and it will send verification email to you email address provided in above step.
6. Check your email inbox and click on the password setup link. It will verify your email address and redirect to password set up page from where you can set your password.


# Free Community Edition

Free community edition best suits for educational users and students.

DatabeanStalk invite educational users and students to sign up free community workspace for educational needs and proof of concepts.


# Sign up

<mark style="color:blue;">**Sign up for Free Community Edition**</mark>

1. Click on Sign up and select Free Community Edition.
2. Enter you name, email, company, contact number and click on Get Started.
3. It will send verification email to your email accress.
4. You can open email inbox and click on verification link and set up password.
5. [Login](https://apps.databeanstalk.com/Identity/Account/Login?ReturnUrl=%2F) to your Free Community Edition version.

![](/files/RrIDeHCLjWs78F6nzNva)


# Login

<mark style="color:blue;">**Login into Free Community Edition**</mark>

After Password Setup you will be redirected to the login page of DatabeanStalk. You can also open [login](https://community.databeanstalk.com/dashboard/) page from DatabeanStalk website header.

<mark style="color:blue;">**DatabeanStalk Product Landing Page**</mark>

Now you will be in DatabeanStalk product login page first. Here you have to enter same credentials once again.

![](/files/5VIf4rCmuSyK2bXF31Zl)

After Successful login it will redirect to DatabenStalk product landing page.

![](/files/S1JollceKDjwLb5wGPCx)


# Data Science Notebook

### JupyterHub

Click on JupyterHub from left menu or click on QuickStart Jupyterhub, it will open up new secure window with your same user credentials

![](/files/mOuHW45MdM6UpbFs92gu)

### Server Options

Select server options for your Data Science run time environment with number of CPU core and RAM and click on start.

![](/files/fCrIHfCghRbJauM5wlGn)

### Data Science Runtime Environment

Start Data science Python environment from JupyterHub launcher and open sample notebook for quick test.

![](/files/KtKD84gyZrUxbwDJwuEi)

Databeanstalk pre-load many example use case notebooks for you in your home directory.


# Data Analysis using Spark

In this tutorial, you'll learn the basic steps to load and analyze data with Apache Spark in DatabeanStalk PySpark environment.

### Apache Spark Runtime Environment

Click on JupyterHub from left menu or click on QuickStart Spark or PySpark, it will open up new secure window with your same user credentials

![](/files/mOuHW45MdM6UpbFs92gu)

### Server Options

Select server options for **Run Managed Apache Spark environment** and click on start.

![](/files/fCrIHfCghRbJauM5wlGn)

Jupyter notebook start with multiple Spark runtime kernels like native python, PySpark, R or core Spark.

![](/files/WNYqAjZj2xt7HvtFIspk)

Initialize spark session with "sc" in notebook and Databeanstalk create Spark driver and two executors in kubernetes Databeanstalk plateform.

![](/files/QJMmI6A48IarRiAdkjEx)


# Projects

{% hint style="info" %}
**Good to know:** Splitting your product into fundamental concepts, objects, or areas can be a great way to let readers deep dive into the concepts that matter most to them.
{% endhint %}


# Members

{% hint style="info" %}
**Good to know:** Splitting your product into fundamental concepts, objects, or areas can be a great way to let readers deep dive into the concepts that matter most to them.
{% endhint %}


# Task Lists

{% hint style="info" %}
**Good to know:** Splitting your product into fundamental concepts, objects, or areas can be a great way to let readers deep dive into the concepts that matter most to them.
{% endhint %}


# Tasks

{% hint style="info" %}
**Good to know:** Splitting your product into fundamental concepts, objects, or areas can be a great way to let readers deep dive into the concepts that matter most to them.
{% endhint %}


# For Designers

{% hint style="info" %}
**Good to know:** depending on the product you're building, it can be useful to explicitly document use cases. Got a product that can be used by a bunch of people in different ways? Maybe consider splitting it out!
{% endhint %}


# Figma Integration

{% hint style="info" %}
**Good to know:** depending on the product you're building, it can be useful to explicitly document use cases. Got a product that can be used by a bunch of people in different ways? Maybe consider splitting it out!
{% endhint %}


# For Engineers

{% hint style="info" %}
**Good to know:** depending on the product you're building, it can be useful to explicitly document use cases. Got a product that can be used by a bunch of people in different ways? Maybe consider splitting it out!
{% endhint %}


# GitHub Integration

{% hint style="info" %}
**Good to know:** depending on the product you're building, it can be useful to explicitly document use cases. Got a product that can be used by a bunch of people in different ways? Maybe consider splitting it out!
{% endhint %}


# For Support

{% hint style="info" %}
**Good to know:** depending on the product you're building, it can be useful to explicitly document use cases. Got a product that can be used by a bunch of people in different ways? Maybe consider splitting it out!
{% endhint %}


# Intercom Integration

{% hint style="info" %}
**Good to know:** depending on the product you're building, it can be useful to explicitly document use cases. Got a product that can be used by a bunch of people in different ways? Maybe consider splitting it out!
{% endhint %}


# Keyboard Shortcuts

{% hint style="info" %}
**Good to know:** depending on the product you're building, it can be useful to explicitly document use cases. Got a product that can be used by a bunch of people in different ways? Maybe consider splitting it out!
{% endhint %}


