แสดงบทความที่มีป้ายกำกับ regression แสดงบทความทั้งหมด
แสดงบทความที่มีป้ายกำกับ regression แสดงบทความทั้งหมด

วันจันทร์ที่ 13 ธันวาคม พ.ศ. 2553

logistic regression analysis - understanding odds and probability


Image : http://www.flickr.com


to measure the probability and the same probability: the probability of a given result. People use the terms interchangeably possibilities and the chances of a casual use, but this is regrettable. It only creates confusion, because they are not equivalent. They measure the same thing on different scales. Imagine how confusing it would be if people used interchangeably Celsius and Fahrenheit. "There will be 35 degrees today" might actually wear the wrong way.

Remember meback to your introductory course in statistics back to all these problems on the probability of drawing red balls and white balls from an urn. In these problems, the probability of drawing a red ball is measured by how many balls there were in total and how many were red.

In measuring the probability of a result, we need to know two things: how many times something happened and how often it could happen. The result of interest is a success, if it is a good result orno.

The other exit is a failure. Every time you encounter the results is a process called. Since each process in the success or failure, the number of successes and failures in the number of total order must be based on the total number of attempts.

Probability of success is the number the total number of attempts has occurred with respect.

Chances are the number of successes has been the number of errors occurred in the comparison.

For example, the likelihood of accidents in a forecastparticular intersection, every vehicle that is going through an intersection as an attempt. Each study is one of two results: pass or accident. If the result we are most interested in the modeling of an incident that happened (no matter how it sounds morbid) is.

Probability (success) = number of successes / total number of units attempted (success) = number of successes / number of failures

The odds are often written as:

Number of successes: 1 failures

Read 'Number of hits for all faults 1. But often, one will be deleted.

I see a lot of learning when the researchers blocked logistic regression because they are not on the scale used probability thinking of a bet.

Equal opportunities are a first success for every failure 1. 01:01 equal probability .5. A success for all the 2 studies.

The odds are infinity to 0. Odds greater than 1 indicates success rather than failure. Rates of less than 1 indicates the failure is moreas a success.

Probability can range from 0 to 1-area. probability greater than 0.5 indicates success rather than failure. less than 0.5 indicates an error probability is more likely to be a success.

Example: In the last month, shows data from a particular intersection, a .354 that the car drove by him, 72 it was an accident.

72, 1282 incident = = Errors Safe Passage (1354-1372) Total Error = - Pr happened (accident) = 72/1354 = 0.053 Pr Safe (Passage) = 1282/1354 = 0.947 Odds (accident) = 72/1282 = 0.056 Odds (Security) = 1282/72 = 17.87

Now you get the computer because you will see how these relate to each other.

Odds (accident) = Pr (accident) / Pr (security) Odds (accident) = (72/1354) / (1282/1354) = 0.056 (the denominator cancel) Odds (accident) = 1/Odds (Safe Passage) = 1/17.87

วันอังคารที่ 7 ธันวาคม พ.ศ. 2553

Linear Regression Analysis - When NOT to Center a Continuous Predictor Variable


Image : http://www.flickr.com


There are two reasons to center predictor variables in any time of regression analysis - linear, logistic, multilevel, etc.

1. To lessen the correlation between a multiplicative term (interaction or polynomial term) and its component variables (the ones that were multiplied).
2. To make interpretation of parameter estimates easier.

But when is centering NOT a good idea?

Well, basically when it doesn't help.

For reason #1, it will only help if you have multiplicative terms in a model. If you don't have any multiplicative terms - no interactions or polynomials - centering isn't going to help.

For reason #2, centering especially helps interpretation of parameter estimates (coefficients) when:

a) you have an interaction in the model
b) particularly if that interaction includes a continuous and a dummy coded categorical variable and
c) if the continuous variable does not contain a meaningful value of 0
d) even if 0 is a real value, if there is another more meaningful value, such as a threshhold point. (For example, if you're doing a study on the amount of time parents work, with a predictor of Age of Youngest Child, an Age of 0 is meaningful and will be in the data set, but centering at 5, when kids enter school, might be more meaningful).

So when NOT to center:

1. If all continuous predictors have a meaningful value of 0.
2. If you have no interaction terms involving any continuous predictors with categorical ones.
3. And if there are no values that are particularly meaningful.

All three of these criteria should apply before you choose to not center. If any one is false, centering will help you interpret your coefficients.

วันอาทิตย์ที่ 26 กันยายน พ.ศ. 2553

Multinomial logistic regression models and ordinal variables


Image : http://www.flickr.com


The multinomial (aka polytomous) logistic regression is a simple extension of binomial logistic regression model. They are used when the dependent variable is greater than two (disordered) has rated categories.

Dummy coding of independent variables is quite common. In multinomial logistic regression, the dependent variable is dummy-coded variables 1 / 0 is a variable for all categories except one, so if there are categories M will be M-1 dummy variables. All except their own category dummy variable. Each category is dummy variable has a value of 1 in its category and a 0 for all others. A category, the category of reference, is not its own dummy variable, as is clearly indicated by all other variables equal to 0.

Logistic regression mulitnomial then estimated a binary logistic regression model separately for each of these dummy variables. The result is M-1 binary> Logistic regression models. Each tells the effect of predictors on the probability of success in this category compared to the reference category. Each model has its own intercept and regression coefficients - the predictors may be of interest to each class may vary.

Why not just run a series of binary regression models? They could, and people were once, in multinomial regression models in the software away. You will probably get similar results. But it workstogether means that they simultaneously estimate the parameter estimates are more efficient means - there are fewer errors, unexplained.

Ordinal Logistic Regression: Proportional Odds Model

If the response categories are ordered, you could have a multinomial regression model. The disadvantage is that you throw away the information about the order. An ordinal logistic regression model retains this information, it is moreare involved.

In the proportional odds model, where the event is not modeled with a score in a single category, as in the binary and multinomial models. Rather, the event is modeled with a score in a particular category or all of the previous category.

For example, for a response variable with three ordered categories, the possible events, defined as:

* In Group 1
* In Group 2, or 1
* In Group 3, 2 or 1

In proportionalQuote model has its own intercept any results, but the same regression coefficients. This means:

1. the overall rate of all cases, be different, but according to the effect of predictors on the probability of an event in any other category, for each category. This is a hypothesis of the model, you need to check. It is often violated.

The model is a bit 'different than usual, written in SPSS, with a minus sign in all the intercepts and regressionCoefficients. This is a convention to ensure that lead to positive coefficients, the increase in X to an increased likelihood of a greater number of response values categories. In SAS, the character is a plus, and elevations of a predictor for increased risk of lead lower numbered response categories. Make sure you understand how the model in your statistics package before interpreting the results.

วันอาทิตย์ที่ 29 สิงหาคม พ.ศ. 2553

The linear regression analysis - three reasons why researchers in psychology, regression, ANOVA should know


Image : http://www.flickr.com


Back when I was doing research in psychology, I knew ANOVA. I had a series of courses that is done, and you could run back and forth. I heard ANCOVA, ANOVA, but in any class that the last argument was about the curriculum, and we are always short of time.

The other thing that drove me crazy was the statistics professors always said, "ANOVA is a special case of regression." I could not for the life of me understand how or why.

Not beforeI turned on the statistics, I finally found a class of regression and found that everything was ANCOVA. And only when I started consulting, and see hundreds of different models of regression and ANOVA finally got the connection.

But if you're not driving curiosity ANOVA and regression, regression want to know why you should, as a researcher in psychology, education or agriculture, which forms in ANOVA? There are three main reasons.

There is a firstmany, many variables and continuous independent covariates should be included in the models. Without the tools to analyze as continuous, have left, forcing the shares ANOVA with any technology, as the median. At best, it is losing power. At worst, you are publishing your article because you are missing a real impact.

According After a solid understanding of the general linear model in its various forms equipped to truly understand the variablesand their relationships. It allows to model a different way - I try not fishing data, but in uncovering the true nature of the relationship. Having the ability to interact a term or a term add to the square allows you to hear your data and makes you a better researcher.

The third multiple linear regression model is the basis for many other statistical methods - logistic regression, multilevel models and mixed Poisson regression, survival analysis,and so on. Each of these represents a step (or smaller jump) on multiple regression. If you have problems with what variables or interactions at the center means to interpret the learning of a These other techniques is difficult, if not painful.

Have thousands of researchers led to their statistical analysis of the last 10 years, I am convinced that a strong intuitive understanding of the general linear model in its variety of forms, the key is not only asafe and effective statistical analyst. You are then free to learn and explore other methods if necessary.

วันจันทร์ที่ 2 สิงหาคม พ.ศ. 2553

Machine learning data - logistic regression with adjustment L2 Python


Image : http://www.flickr.com


logistic regression

Logistic regression is used for binary classification problems - where you have some examples of "on" and other examples that are "outside". It takes as input a career that said, some examples of each class with a label, where each example is "on" or "off." The goal is to learn a model from training data, so that the caption of the new examples that you have not seen and can not predictKnowing the label.

For an example: Suppose you have a lot of data of buildings and earthquakes (for example, the year the building was constructed to describe the type of material used, the strength of the earthquake, etc.), and you know where every building collapsed (ON) or not ("off") in each of the last earthquake. Using these data, you want to make predictions about whether a particular building will collapse in a hypothetical future earthquake.

One of the first models, which would be worthattempts, logistic regression.

Encodes the

He was not working on this exact problem, but I had to quit a job. In practice what they preach, I started looking for a dead simple Python class logistic regression. The only requirement is that I wanted to support the legalization L2 (more on that later). They are also code-share with a group of other people on many platforms, so I wanted as few dependenciesexternal libraries as possible.

I have not found exactly what I wanted, so I decided to take a walk in the past and I use. I've written in C + + and Matlab, but never before in Python.

I'm not the discharge, but there are many good explanations out there to follow if not a bit afraid of calculation '. Just do a little 'Googling for "derivation of logistic regression." The idea is to write the probability of dataas some internal settings of the parameter, the derivative, which show how to modify the internal parameters to make the data more likely. Got it? Good.

For those of you out there who know, inside and outside of logistic regression, see how short the train () method. I like how easy it is to do in Python.

Regularization

I caught a bit 'indirect flak speak during March Madness season, asI settled into my latent carriers of the matrix factorization model of team offensive and defensive strengths in predicting the outcome of NCAA basketball. Apparently, people thought I was stupid - crazy, right?

But seriously, people - legalization is a good idea.

I would home the point. Check out the results of running the code (below) is connected.

Take a look at the top row.

On the left is set training. There are25 examples from the x-axis position and the y-axis indicates whether the sample is "on" (1) or "off" (0). For each of these examples, there is a vector that describes its attributes, which I understand. After training the model, ask the training model that is developed to bypass the labels and the likelihood that each label is "on" only on the basis of the description and examples of carriers has learned that the model (estimate we hope things like strongest earthquake and old buildingsincrease the probability of collapse). Chances are shown red Xs. Top left, the red X on the right are the top of blue dots, it is very safe on the labels of the examples, and that is always correct.

Now, on the right side we have some new examples that the model has never seen before. This is called the test in September This is essentially the same as the left, but knows nothing of the test model of class labels (yellow dots). WhatYou see, there is still a decent job of providing the label, but there are some cases where it is worrying very confident and very wrong. This is known as overfitting.

This is where regularization a. While walking between the lines about, we will be stronger L2 regularization - or, equivalently, pressure on the internal parameters to zero. This has the effect of reducing the model of certainty. Just because it's perfectly reconstruct training setdoes not mean that you have discovered everything. You can imagine that if you rely on this model to make critical decisions, it would be desirable to have at least a little 'there in the regularization.

And here is the code. Seems long, but most of it is to generate the data, then the result. The bulk of the work is done by train () method, which only three (thick) lines. It requires NumPy, SciPy and pylab.

* For full disclosure, II admit that generates random data in order so that it is vulnerable to overfitting, logistic regression, without looking at regularization perhaps worst of them.

Python code

Importing scipy.optimize.optimize fmin_cg, fmin_bfgs, fmin

Import NumPy as NP

final sigma (x):

Return 1.0 / (1.0 + np.exp (-x))

Class SyntheticClassifierData ():

def __init__ (self,N, D)

"" "Create instances of input vectors and N d-dimensional 1D

Class labels (-1 or 1). ""

Mean = 0.05 np.random.randn * (2, d)

np.zeros self.X_train = ((Nd))

np.zeros self.Y_train = (N)

for i in range (N):

if np.random.random ()> 0.5

y = 1

Other:

y = 0

self.X_train [i:] = np.random.random (d) + y medium [:]

self.Y_train [i] = 2.0 * y - 1

self.X_test np.zeros = ((Nd))

self.Y_test np.zeros = (N)

for i in range (N):

if np.random.randn ()> 0.5

y = 1

Other:

y = 0

self.X_test [i:] = np.random.random (d) + y medium [:]

self.Y_test [i] = 2.0 * y - 1

Class LogisticRegression ():

"" "A simple logistic regression models L2 regularization(Zero-mean

priori Gaussian parameters). ""

def __init__ (self, x_train = None, y_train = None, x_test = None, y_test = None,

alpha =. 1, summary = False):

# Set the strength of regularization L2

self.alpha =Alpha

# Set the data.

self.set_data (x_train, y_train, x_test, y_test)

# Initialize the parameters to zero in the absence of a better choice.

self.betas np.zeros = (self.x_train.shape [1])

DEFnegative_lik (self, beta):

return -1 * self.lik (beta)

def lik (self, beta):

"Probability" of data according to current settings of the parameters. "

# A probability

L = 0

forself.n in range ():

+ L = log (sigma (self.y_train [i] *

np.dot (Beta self.x_train [i ,:])))

# BeforeProbability

for k in range (1, self.x_train.shape [1]):

L -= (self.alpha / 2.0) * self.betas [k] ** 2

Back to the

def train (self):

"" "Define the slope and hand out a SciPyGradient-based

Optimizer. ""

# Definition of the derivative of probability than beta_k.

# It is necessary to multiply by -1, because there will be minimal.

dB_k lambda = B, K: np.sum ([* [B-self.alpha k]+

self.y_train [i] * self.x_train [i, k] *

Sigma (-self.y_train [i] *

np.dot (Bself.x_train [i ,:]))

for i in (self.n range)]) * -1

# The full course is a series of derivatives componentwise

dB = lambda B: Np.array ([dB_k (B, K)

for k in range (self.x_train.shape [1])])

Optimize #

self.betas fmin_bfgs = (self.negative_lik, self.betas,fprime = dB)

final set_data (self, x_train, y_train, x_test y_test):

"" Take the data that has already generated. " ""

self.x_train = x_train

self.y_train = y_train

self.x_test = x_test

self.y_test = y_test

y_train.shape self.n = [0]

final training_reconstruction (self):

p_y1 np.zeros = (self.n)

for i in range (self.n):

[I] = p_y1 sigmoid (np.dot (self.betas,self.x_train [i ,:]))

Back p_y1

test_predictions DEF (self):

p_y1 np.zeros = (self.n)

for i in range (self.n):

[I] = p_y1 sigmoid (np.dot (self.betas,self.x_test [i ,:]))

Back p_y1

plot_training_reconstruction final (self):

plot (np.arange (self.n) self.y_train, 0.5 + 0.5 * "Bo")

plot (np.arange (self.n) self.training_reconstruction ()'RX')

ylim ([-. 1, 1.1])

plot_test_predictions DEF (self):

plot (np.arange (self.n) 0.5 + 0.5 * self.y_test, 'yo')

plot (np.arange (self.n) self.test_predictions (), 'RX')

ylim ([-. 1, 1.1])

if __name__ == "__main__":

Import pylab*

# 20 dimensional data set to create with 25 points - this is

# Sensitive to overfitting.

data SyntheticClassifierData = (25, 20)

# Run for a variety of strengths regularization

Alpha = [0, .001, .01, 0.1]

for j, a in enumerate (Alpha):

# Create a newLearners, but use the same data for each cycle

LR = LogisticRegression (= x_train data.Y_train data.X_train, y_train =

x_test = data.X_test,y_test = data.Y_test,

alpha = a)

print "initial probability:

printlr.lik (lr.betas)

# Train model

(Lr.train)

# Display version more

print "Final Beta"

Print lr.betas

print "Final lik"

Print lr.lik (lr.betas)

# Plot the results

Subplot (len (alpha), 2, 2 * j + 1)

lr.plot_training_reconstruction ()

ylabel ('alpha ='% s% a)

if j== 0:

Title ("reconstructions Student Set)

Subplot (len (alpha), 2, 2 * j + 2)

lr.plot_test_predictions ()

if j == 0:

Title ("Test SetForecast)

show ()