[Get it solved] Use the Cluster Analysis and Decision Tree induction algo...

Check Out Our Work & Get Yours Done

Submit Work

Download Sample

Enroll in the complete course for only $250 USD*

Order Now

Submit work Offers

Use the Cluster Analysis and Decision Tree induction algorithm on a weather forecast problem. It is a binary classification problem to predict whether or not a location will get rain the next day.

statistics

Description

Use the Cluster Analysis and Decision Tree induction algorithm

on a weather forecast problem. It is a binary classification problem to predict whether or not a location will get rain the next day.

Information about the dataset (Weather Forecast Training.csv):

• Location: The location name of the weather station

• MinTemp: The minimum temperature in degrees celsius

• MaxTemp: The maximum temperature in degrees celsius

• Rainfall: The amount of rainfall recorded for the day in mm

• Evaporation: The so-called Class A pan evaporation (mm) in the 24 hours to 9am

• Sunshine: The number of hours of bright sunshine in the day.

• WindGustDir: The direction of the strongest wind gust in the 24 hours to midnight

• WindGustSpeed: The speed (km/h) of the strongest wind gust in the 24 hours to midnight

• WindDir: Direction of the wind

• WindSpeed: Wind speed (km/hr) averaged over 10 minutes

• Humidity: Humidity (percent)

• Pressure: Atmospheric pressure (hpa) reduced to mean sea level

• Cloud: Fraction of sky obscured by cloud This is measured in “oktas”, which are a unit of eigths. It

records how many eigths of the sky are obscured by the cloud. A 0 measure indicates a completely clear sky

whilst an 8 indicates that it is completely overcast.

• Temp: Temperature (degrees C)

• RainTodayBoolean: 1 if precipitation (mm) in the 24 hours to 9 am exceeds 1mm, otherwise 0

• RainTomorrow: The target variable. Did it rain tomorrow?

Organize your report using the following template (with the section breakdown and grading rubrics):

Section 1: Data preparation (30%)

• Discuss the potential data quality issues you identify about the dataset and how you apply various data preprocessing techniques to cope with those issues and perform Exploratory Data Analysis (EDA). Specifically, discuss the type of techniques you carry out in order to prepare the dataset for the machine learning algorithms you use in the next section. Whenever appropriate, enhance your EDA with effective data visualization.

Section 2: Build, tune and evaluate cluster analysis and decision tree models (50%)

• Apply both clustering algorithm (kmeans and HAC) and decision tree induction algorithm to the weather forest training data and construct models. Perform extensive model experiments with hyper-parameters’ tuning. Discuss your choice of hyper-parameters for each algorithm and produce tables summarizing the best performing models and their corresponding model specifications (i.e. the combination of hyper-parameters). Also, explain your choice of model performance evaluation methods and metrics in order to produce unbiased and low variance estimates.

• In your analysis writeup, include the discussion regarding how to repurpose the unsupervised learning algorithms like clustering for classification and how to judge the performance of the algorithms.

• For decision tree induction algorithm model performance evaluation, generate Receiver Operating

Characteristics (ROC) curve and calculate the Area Under Curve (AUC) metric for the identified best performing model. Include a decision tree output visualization and interpret the decision tree model.

• Include a detailed explanation of your modeling process and interpretation of the results in your analysis

writeup (with markdown language) and structure such writeup in an easy-to-follow layout. Please limit the program output only to the most relevant part which is used to support your analysis. An excessive amount of less relevant outputs (e.g. display the whole dataset) in your report is not needed

Section 3: Prediction and interpretation (20%)

• After building the classification models, apply them to the test dataset (Weather Forcast Testing.csv)

provided to predict if each location will rain tomorrow.

•Needed prediction results as a CSV file with four columns (ID, kmeans, HAC, DT) for the

classification results out of the pre-specified machine learning algorithms respectively.

Needed rmarkdown or notebook document together with a knitted report (HTML format) and prediction CSV file.

Related Questions in statistics category

The Central Dogma of molecular biology

The assignment requires construction, development and interpretation of a model to solve a business problem.

The deliverable is a Word document with your answers to the questions posed below based on the data you find. Required Software

You will read an Excel worksheet into SPSS and solve a problem

In this assignment, you will evaluate two potential data collection methods for your qualitative research study.

A DNP student has finished collecting the data for his/her DNP quality improvement project.

Coca-Cola Company is the biggest independent bottler of Coca-Cola, located in the United States. The company is famous for its flagship product, Coca-Cola, among other beverage products.

This Individual Coursework involves investigating properties of orthogonal matrices, the bivariate normal distribution, and investigating some of the links between them using small examples involving numpy and matplotlib

Our group is conducting a study for our class, Statistics of Sociology 3307. For our final project, we are examining how aware first-generation students are of resources at the Cal Poly Pomona.

WM volumes of the SCG will be the continues variable. Due to potential differences between the WM volumes of men and women, the effect of gender will be assessed.

Get Higher Grades Now

Tutors Online

Description

Drop Files Here Or Click to Upload

April

January

February

March

April

May

June

July

August

September

October

November

December

2025

1950

1951

1952

1953

1954

1955

1956

1957

1958

1959

1960

1961

1962

1963

1964

1965

1966

1967

1968

1969

1970

1971

1972

1973

1974

1975

1976

1977

1978

1979

1980

1981

1982

1983

1984

1985

1986

1987

1988

1989

1990

1991

1992

1993

1994

1995

1996

1997

1998

1999

2000

2001

2002

2003

2004

2005

2006

2007

2008

2009

2010

2011

2012

2013

2014

2015

2016

2017

2018

2019

2020

2021

2022

2023

2024

2025

2026

2027

2028

2029

2030

2031

2032

2033

2034

2035

2036

2037

2038

2039

2040

2041

2042

2043

2044

2045

2046

2047

2048

2049

2050

Sun	Mon	Tue	Wed	Thu	Fri	Sat
30	31	1	2	3	4	5
6	7	8	9	10	11	12
13	14	15	16	17	18	19
20	21	22	23	24	25	26
27	28	29	30	1	2	3

00:00

00:30

01:00

01:30

02:00

02:30

03:00

03:30

04:00

04:30

05:00

05:30

06:00

06:30

07:00

07:30

08:00

08:30

09:00

09:30

10:00

10:30

11:00

11:30

12:00

12:30

13:00

13:30

14:00

14:30

15:00

15:30

16:00

16:30

17:00

17:30

18:00

18:30

19:00

19:30

20:00

20:30

21:00

21:30

22:00

22:30

23:00

23:30

Get Free Quote!

357 Experts Online

Get Instant Help with your Questions &
boost your grades

you can count us with it
Highly Satisfied Students 4.9/5
Based On 19835+ Reviews

Get Help Now

We Provide Services Across The Globe

Disclaimer: The reference papers or solutions provided by Calltutors.com serve as model papers or solutions for students or professionals and are not to be submitted as it is to any institutions. These documents are intended to be used for research and reference purposes only. University and company's logo's are the property of respected owners. We don't have affiliation with the mentioned universities. By using our services means, you agree to our Honor Code , Privacy Policy , Terms & Conditions , Payment , Refund & Cancellation Policy.

Enroll in the complete course for only $250 USD*

Use the Cluster Analysis and Decision Tree induction algorithm on a weather forecast problem. It is a binary classification problem to predict whether or not a location will get rain the next day.

statistics

Description

Get instant assignment help service

Related Questions in statistics category

Policy

Exploring

Other

Connect With Us

Get Instant Help with your Questions &
boost your grades

you can count us with it
Highly Satisfied Students 4.9/5
Based On 19835+ Reviews

We Provide Services Across The Globe

Enroll in the complete course for only $250 USD*

Use the Cluster Analysis and Decision Tree induction algorithm on a weather forecast problem. It is a binary classification problem to predict whether or not a location will get rain the next day.

statistics

Description

Get instant assignment help service

Related Questions in statistics category

Policy

Exploring

Other

Connect With Us

Get Instant Help with your Questions & boost your grades

you can count us with it Highly Satisfied Students 4.9/5 Based On 19835+ Reviews

We Provide Services Across The Globe

Get Instant Help with your Questions &
boost your grades

you can count us with it
Highly Satisfied Students 4.9/5
Based On 19835+ Reviews