레이블이 Statistical Data Mining인 게시물을 표시합니다. 모든 게시물 표시
레이블이 Statistical Data Mining인 게시물을 표시합니다. 모든 게시물 표시

2014년 1월 23일 목요일

Statistical Data Mining - FInd the Mr / Miss Right

Amy Webb, the speaker on that video, gave a speech about how she met her husband through online dating.

How she had done is similar to the way that a mathematician used. This is a link about how a mathematician found her girl friend. Link
The fabulous a data mining can give to the data miner is that once you have desire to know or find out something, you can gather, process and modify data to achieve your goal through data mining.

2014년 1월 20일 월요일

Curriculum in Statistics

2014년부터 통계학 연구실에 진학할 학생으로써, 통계학 커리큘럼을 확인하고자 한다. 현재 내가 수강한 과목들은 <확률 및 통계> <수리통계학> 뿐이다. 과목은 1학년부터 순서대로 적어본다. (서울대는 봄학기만 볼 수 있어서, 봄학기만 적어보았다.)

서울대

  1. 1학년
    • 통계학 - KAIST의 확률통계 과정. ANOVA 까지 다룬다. 교재는 일반통계학 - 김우철
    • 통계학실험
  2. 2학년
    • 통계학의 개념 및 실습 : 통계학과 비슷한 과목으로 보인다.
    • 확률의 개념 및 응용 : 기초확률론과 확률통계의 사이에 위치한 과목으로 보인다.
  3. 3학년
  4. - 계량경제학에서 반드시 커버할 것.
    • 수리통계
    • 회귀분석 및 실습 - 교재-Introduction to Linear Regression Analysis-Montgomery, Peck & Vining-Wiley-2012
    • 실험계획법
  5. 4학년
  6. - 계량경제학에서 반드시 커버할 것
    • 베이즈통계 및 실습
    • 비모수통계 및 실습
    • 시계열분석 및 실습

고려대 정경대학 통계학부

  1. 1학년
    • 통계적탐구
    • 기초통계 워크샵
  2. 2학년
    • 통계수학
    • 행렬이론 - 선형대수학
    • 통계프로그래밍 입문
    • 기초확률론
    • 회귀분석
    • 사회과학을 위한 통계적 방법
  3. 3학년
  4. - 계량경제학에서 반드시 커버할 것.
    • 수리통계
    • Nonparametric method
    • 실험계획법
  5. 4학년
  6. - 계량경제학에서 반드시 커버할 것
    • 다변량통계분석
    • 응용통계 - 금융통계, Biostatistics
    • Data mining 입문

다음은, 내가 수강하지 않은 과목 위주로, CMU Curriculum 을 적어본다.

CMU Department of Statistics

  • Prerequisite
    1. Statistical foundation - Regression
    2. Experiment design
  • Disciplinary Core
    1. 경제학 - 미시, 거시, 국제, 계량
    2. 기초확률론
    3. 회귀분석의 응용
    4. Data analysis

즉, 내가 수강하거나 공부해야 하는 과목들은, 다음과 같다.
  1. 경제학계열 : 계량경제학
  2. 통계학계열 : Regression analysis, Experiment Design, 고급통계학
  3. 확률론계열 : 확률론, 확률과정론

2014년 1월 18일 토요일

What is 'R'?

2014년 1월~2월, 학부 졸업과 석사 입학의 징검다리에 있는 입장이다. 이 시간동안 어차피 공부도 잘 되지 않고, 무엇을 배워두면 참 좋을까 고민하던 차에, 번뜩 생각난 아이디어. "R을 배워보자!"

2013년 방학 등 시간이 날때마다 틈틈히 C++, Python, MATLAB 및 프로그래밍 언어는 아니지만 HTML, LaTex를 익혀온 까닭에, 프로그래밍은 이제 자신이 있다. 이 새로운 언어가 나의 연구에 얼마만큼 편의를 제공해 줄 수 있을지 기대가 된다.

이 언어에 어떻게 접근해야 할까? 내가 구한 몇 가지 튜토리얼들의 목차를 정리하면 다음과 같다. (R 다운로드, 설치 등 당연한 정보들은 건너뛴다.)

R cookbook

  1. 변수 설정해보기
    • 변수설정하기 (vector 등)
    • 함수정의하기
  2. R에서 제공하는 기본기능들
    • Command history 보기
    • Script 돌려보기
  3. 입출력
    • 직접 데이터 입력하기
    • 입력된 데이터 반올림하기
    • 파일로 직접 출력해보기
    • CSV 파일 읽어오기
    • HTML 등 웹페이지에서 직접 데이터 긁어오기
  4. 데이터구조
    • 데이터를 벡터로 바꾸기
    • Matrix 다루기
    • String data 다루기
  5. 확률 & 통계
    • Combination, Permutation 계산
    • Random generating
    • Calc. Prob.
    • Quantile
    • Quantile
    • 회귀분석 및 ANOVA
    • 시계열분석


R을 이용한 통계프로그래밍

  1. 기본 입출력 명령어
  2. R object
    1. 데이터의 종류
    2. 벡터
    3. array & matrix
    4. list
    5. data frame - 우리가 생각하는 표
  3. 데이터 읽어오기
  4. R 프로그래밍
  5. R과 확률통계

2014년 1월 15일 수요일

Introduction of Statistical Data Mining

Statistical Data Mining
본문은 위의 글을 참조 / 번역하고, 필자가 주석을 단 글입니다.

Tutorial

Classification Algorithm, Regression Algorithm, Data Mining Operation 으로 나뉜다.
  • Decision Tree 가장 많이 쓰이는 Classification 기법. Information gain 이 추후에 들어오는 데이터를 어떻게 잘 모형화할 수 있는지 설명한다.
  • Information gain Entropy 이론을 다룬다. Entropy 는 Information gain의 가장 중요한 Measure 로 활용된다.
  • Probability 기본적인 확률지식을 다룬 이후에, Density estimation 등을 다룬다. 그 후에 Bayes 통계방법론으로 연결된다. 마지막으로, Multivariate density function 으로 연결되는 모형이다.
  • Gaussian 검색필요
  • MLE Parameter 를 찾는 Technique를 다룬다.
  • Cross Validation 기존의 Data를 바탕으로 Model 을 구축했을 때, future unseen data 를 얼마나 잘 설명할 지 말해주는 '설명력'에 관한 Topic 이다.
  • Neutral Networks 먼저 Linear Regression 부터 시작한다. 이를 통해 SSE 방법을 도출한다. 이 외에 Nonlinear Model 에 대해서도 다룬다.
  • Regression Algorithm Regression Trees, Cascade Correlation, Group Method Data Handling (GMDH), Multivariate Adaptive Regression Splines (MARS), Multilinear Interpolation, Radial Basis Functions, Robust Regression, Cascade Correlation + Projection Pursuit (뭐지 모르겠다.)
  • Bayesian Networks 확률모형을 다루고, Joint Distribution 을 다루며, 그것의 Drawback 을 다룬다. 그에 대한 대안으로 Bayesian Statistics 를 소개한다. 이를 이용한 Statistical inference 를 다루기도 한다. A typical use of inference is "I've got a temperature of 101, I'm a 37-year-old Male and my tongue feels kind of funny but I have no headache. What's the chance that I've got bubonic plague?".
  • Gaussian Mixture Model Density Estimation 을 비롯한, Clustering 에 가장 많이 쓰이는 분야이다. Clustering분야를 설명하고, Expectation Maximization 에 대해서 설명한다.
  • Markov Model DTMC, CTMC
  • VC dimension Machine learning 의 기초를 다룬다.
  • Game Theory Zero-sum Game Theory 를 다룬다. Non zero game theory 를 다룬다.


The elements of statistical learning - Data mining, Inference, and Prediction by Prof. Trevor Hastie, Robert Tibshirani, Jerome Friedman


Supervised learning

Supervised learning 에서는, Input information을 바탕으로 결과를 예측하는 방법을 학습한다.
  1. Overview of supervised learning
    • Two simple approaches to prediction : Least squares & nearest neighbors
    • Statistical decision theory
    • Statistical Model : Joint distributions & Function approximation
  2. Linear model for regression
    1. Linear regression and Least squares
    2. Shrinkage method
  3. Linear method for classification
  4. Basis Expansion and Regularization
  5. Kernel smoothing method
  6. Model Assessment and Selection
  7. Model inference and averaging
  8. Additive models, Tree, and related method
  9. Boosting and additive trees
  10. neural networks
  11. Prototype method and Nearest Neighbors

Unsupervised learning

Unsupervised learning 에서는, 결과를 예측하지 않는다. 대신, Input measure의 패턴과 관계를 파악하는 방법을 학습한다.
    • Association rules
    • Cluster analysis
    • Principal components, curves and surfaces
    • Matrix factorization
  1. Random forest
  2. Ensemble Learning
  3. Graphical method