Bootstrap Tape 1
Sign in to track this film in your collection or want list.
Description: Title Bootstrap methods American Statistical Association. Author Bradley Efron Robert Tibshirani American Statistical Association American Statistical Association. Annual Meeting (1985 : Las Vegas, Nev.) Subjects Bootstrap (Statistics) Publisher Alexandria, VA : American Statistical Association ; Cincinnati, Ohio : The Film House Creation Date 1985 Format 8 videocassettes sd., col. ; 3/4 in.. Notes "FH 382"--Container.; Title from cass. label.; "Examples"--Lecture 4.; Videorecording of a short course held during the1985 annual meeting of the American Statistical Association. Performers Bradley Efron, Robert Tibshirani. Language English Record id 01IASU_ALMA21225953260002756 MMS ID:990016974940102756 Source 01IASU_ALMA
Transcription
hello can you hear me I see one hit how about in the back people in the back of the room hear me raise your hand if you can hear me thank you thank you good morning I'm Tom Devlin and on behalf of the short-course Committee of the American Statistical Association it is my pleasure to welcome you to Las Vegas and to our annual short course program this year the continuing education office is pleased to present two short courses on computer intensive methods today grad Efren and Rob chip shirani will introduce us to the bootstrap and tomorrow jerome Friedman Richard Olson and Charles Stone will discuss classification and regression trees most statistical methods in common use today rely on the assumption of normality and focus on statistical measures whose theoretical properties are mathematically tractable powerful high-speed digital computers have stimulated the development of new statistical analysis statistical methods and theories which release researchers from these limitations these methods are computational spendthrifts and so have been referred to as computer intensive one class of such methods is the bootstrap the bootstrap was invented less than 10 years ago by Bradley Efrain the purpose of this course is to provide you with a working knowledge of the bootstrap to aid in this goal we are fortunate to have as our instructors Bradley Efrain and Robert tips irani let me introduce you to them Bradley Efrain his professor of statistics and biostatistics at Stanford University where he is currently chair of the program in mathematical sciences he is a former editor of the theory and methods section of Jassa and he is a fellow of the American Statistical Association of the Institute of mathematical statistics and a recipient of the prestigious MacArthur award The Wall Street Journal recently had a front-page article on these awards it was quite an interesting article indicating that only 166 people have received such an such an award and contain and it contained interviews with a few recipients one of those interviewed had an interesting comment about her fellow recipients she said these are not people who go to Las Vegas Robert tips irani is a postdoc fellow in the department of preventive medicine and biostatistics at the University of Toronto he's a recent PhD graduate of Stanford University where his mentor was badly effort his research interests include the bootstrap and nonparametric regression techniques in particular generalized additive models it is now my pleasure to turn the floor over to Professor Aaron so if I can untagged listen I'm gonna take off my coat okay want to cut down here yeah that's a good right there hey comfortable yeah maybe put this in your back okay thank you well that was supposed to establish that I have a sport coach and that's why we're on we actually have TV recording this for future use and I wanted everyone to see that I indeed own a very nice sport coat but I never wear it for talks because it seems to me to stop necessary motions like why don't you guys ask more questions and that's an important point which I'd like to bring up right at the beginning we're going to be a this is not a one-hour talk we're going to be going for quite a while and it really is deadly if I and Rob drone on and on with no change of voice and presenting material to you the we've given this as a practice course once or twice and it really works a lot better if people will ask questions most of the questions the bootstrap is not a complicated it's not a technically complicated subject so you won't be overwhelmed with technicalities hover at a certain point and I've given lots of bootstrap talks it always occurs to people and to me too why are we doing this and can this possibly be right and those kinds of questions are very helpful if you ask a question just raise your hand I'll say oh yes and then just ask your question in a loud voice I'll repeat the question to make sure that everyone can hear it and then I'll maybe even try and answer the question so that's the way it'll go you have lots of stuff that we gave you you have copies of all our transparencies almost dog we added a couple at the last minute but mostly you have the transparencies that you'll be seeing up there and that's so that you don't have to desperately take know some people feel more comfortable taking notes than not so please do if you feel like it but basically you won't have to take any notes you can just concentrate on thinking of good questions the you also have a report a copy of report called the bootstrap method of assessing statistical accuracy and that's by Rob and me and as a modest goal of of today's presentations you should at the end of the day be able to read the report rather easily and the report has quite a bit more in it of course and will say and gets more into the technical details this is not a good this kind of form is not a good form for getting in deeply into technical stuff and we will try and avoid it other of course some of the things would be I mean it's not English literature or something like that but it will there will be some technical stuff but the the of course the real point of going to something like this is not to learn any particular this or that but to understand the point of view of a different subject matter and we very much hope that we'll get our point of view across to you and the what we want to convince you is that this is a reasonable theory and is in fact the theory that you as applied statisticians or theoretical statisticians know already but in a different guise and what it really is is a version of of the theory that we all use for estimation and setting standard errors brought up to date for the computer age okay so that's about it for getting going of any questions about the rules will go for about an hour or so now then take a break there'll be some coffee then Rob will be up oh yeah we're doing this the way police always have interrogate people in to two-person teams like a good guy and a bad guy and the good guy the bad guy comes in and says something terrible I'm gonna break your arm off something like that and then the good guy comes in and says let me save you from the bad guy confess and Rob's the good guy i-i'll be giving them more basic theoretical stuff of which it's not very theoretical and then Rob will come in and say how this is applied in real real problems and real problems means complicated real problems so we hope that the subject matter of the examples will be of genuine interest to you also ok so let's go the next like this oh I forgot to say one thing well we can leave that up there there's besides the besides the material that I mentioned there's one extra handout you actually got two separate pages that had a lot of numbers on them those were things that didn't show up very well when they got reproduced in the in the handouts and and they're they're the they're the numbers that that are just computer out but for example the first one says on the top bootstrap methods Efrain and tips iran e8 385 now everybody should try and find that piece of paper it'll be on a slide too but you won't be able to read it on the slide so I wanted you to have that piece of paper is there anybody okay wait a minute or two so it says bootstrap methods Efren and tipster on e8 385 and it says at the top and equals 15 observed lifetimes that's going to be it let me put that one over here for a second and then we'll get back to it beyond the main one soon sayin it's not it's not really readable okay is there anybody who has not found that piece of paper then perhaps we can get the people from behind the screen to quit that's a good suggestion thank you I had the terrible fear that the door was already closed okay fine meanwhile I'll handle but I'll handle the slides until ok it's being taken ok so here we are in Las Vegas let last night as we're wandering round Rob said to me here we are in a an entire city built on the strong law of large numbers I thought that was very good comment ok so let's start here we are in a typical situation instantly I work in a biostatistics group and so most of this theory and indeed the bootstrap was generated by the need for certain kinds of data analytic tools so here we are in the typical situation in the stat lab we have some data and I've called the data why and that's just stands for generic data why has in particular this morning we'll be talking about the the situation where Y consists of independent repeated observations a random sample so in I wrote down for example y equals X 1 through X 15 a random sample of 15 life times and here on the the sheet of paper that I said you wouldn't be able to see here's the N equals 15 life time so right at the top of the piece of paper and well I'll put that up back over here in a minute the the fridge I've worn I've ordered the 15 lifetime so the smallest one is 0.14 3 and the next one is 0.18 2 and the biggest one is 2.0 7 6 they're all positive because they're lifetimes they don't they don't look very normal they look like they're sort of stretching out to the right is everybody does anybody not able to see where I'm reading the 15 observed lifetimes ok and typically we'll have a parameter of interest theta that we're interested in knowing for example the true expected lifetime of any one of the typical observations and we'll want to estimate theta from y so that's that's a pretty basic statistical operation and I'm sure that all of you have been in that situation or you wouldn't be here okay and then two basic questions arise the two basic questions are question 1 what statistic theta had of why should we use to estimate theta question 2 how accurate is Theta hat as an estimate of theta so let me repeat the two questions because we're going to come back to these a lot today the first question which you might say is the primary statistical question is how should we estimate what we're interested in how should we use the data to get a number that says something about what we're interested in the second question which is the main one we're going to address today is how accurate is Theta hat having chosen theta hat at the first step how accurate it is it as an estimate of theta and it's our existence as a field depends on the fact that we can answer that question most good scientists can answer the first question rather well without any statistical training that's not exactly true and of course we've had a lot to say about how to answer the first question but typically when doctors or medical researchers come in and talk to me they actually usually have a fairly good idea of how to answer question 1 what question - not only is hard to answer it doesn't even occur to most non statisticians and and yet it's fundamental in some sense to the use of your answer to question one very good scientists people who are trained very well people as good as Linus Pauling if they do not have statistical training are hopeless at answering question 2 in any reasonable way and we're going to be talking about how to answer question 2 next please okay this is to remind you that back in the 1920s a very successful program was begun by RA Fisher to answer both questions and it's I believe the single most successful Enterprise in the history of both theoretical and applied statistics is this program that Fisher developed in just an amazing series of papers that spanned about 1920 through 1935 and the the Fisher's program answered both questions in a most elegant way the answer the question one was used theta-hat of why the MLE that is if you want to answer question 1 what statistic to use you don't have to worry just use the MLE maximum likelihood estimate the answer to that is the answer to question 2 is the how accurate is the MLE was that the standard error of the MLE theta hat is approximately Sigma hat equals 1 over the square root of the Fisher information and I'm sure you've all done those calculations where you differentiate log likelihoods maybe once or twice and take square roots and plug in and you a lot of the answers have become so familiar to us that you don't you forget that that's how they're derived that is you may know the answer ready from from the fact that other people have derived it and written it down in books but in fact an awful lot of statistical theory depends on on these two answers and next please the the bootstrap is a more general way to answer question two and it it involves less parametric modeling than Fisher's theory Fisher's theory is built around always around a parametric model a simple parametric model usually for the for the data in terms of unknown parameters and that's that's a fundamental characteristic of it that we'll talk more about the bootstrap is not necessarily based on parametric modeling it can be used in parametric settings and we'll talk about that it can also be used in nonparametric settings and that's it was originally introduced to be a nonparametric device because it's in the nonparametric settings that Fisher's theory is least satisfactory and least easy to apply and and I must I'll remind you later on that this the bootstrap was not the first attempt to go beyond Fisher's theory there in particular the jackknife which we'll also talk about a little was a very successful and interesting attempt to answer question two for nonparametric situations the bootstrap also works for less smooth statistics I didn't say that this but it the answer to Question 2 for Fisher's theory involves smoothness assumptions you should be able to differentiate the statistics the likelihood functions nicely and all that stuff and the it breaks down for statistics like the median which aren't very smooth at all and we'll talk more about that also so that's the payoff the price is a lot more computation Fisher's theory is extremely elegant and it requires the minimum amount of computation well I actually the right way to say it is it involves the maximum amount of computation you can do on a desk calculator because that is exactly the way Fisher designed the theory he the desk calculator was a gigantic step forward from hand calculation and enabled a theory like Fischer's theory to be to be practical it really wasn't practical without a desk calculator and he many points wrote down that that he wouldn't have invented say analysis of variance which is another part of his theory if there if it hadn't been for the existence of the desk calculator and I remember when I started yeah Stanford our computation room had Monroe calculators and freedom calculators and you sort of decided which one you wanted to use depending on what noise level you could stand because they they they were really like a truck but boy you were glad to get get your hands on one because we're sure a lot better than trying to trying to do multiplications and divisions by hand I went and looked up in the I have a set of 1930 encyclopedia Britannica's that somebody gave me and I like to look up things in them I went up I looked in under computation and the you have to remember that the Britannica at this stage was the supreme arbiter of encyclopedias and presumably had articles written by the best people in the world and the the guy who wrote on computation was not sanguine about the possibility of a mechanical device that actually did division though he thought that a German device that was based on steel balls rolling back and forth had some promise and I like and I like reading these things because first of all if the 1930 encyclopedia is that wrong I have a feeling the 1985 one is not much better we just don't know it yet and and as a reminder the computation has come a long way the bootstrap requires perhaps a hundred or a thousand times more computation than Fisher's theory did and we'll talk about those numbers also but this is quite a bargain because the electronic computers that we now have are perhaps a million times faster and cheaper than those desk Cal calculators and cheap cheaper means it doesn't cost a lot to run if they were a million times faster but it cost a million times more that that wouldn't help any either they really are they really have improved our computational environment by a factor of a million and that's instantly a bigger factor than the step and going from naked-eye astronomy to telescope ii so one can expect that statistics is in for a golden age of of advanced and that lots of theories like bootstrap and like the theory you'll hear about tomorrow will be developed in response to all this power they're the last the first point is the first two points are differences between the two theories the third point though that both theories are completely automatic that is that's a similarity both theories go directly you can actually build a black box that goes that would look something like this let's see if this works here's the black box and here's data and here's say the pyramid the form of the density function let's call it f sub theta the form of the density function and here's the answers I won't use the board much the that is it is possible to write an algorithm for Fisher's method say that will input the data the probabilistic model and output the answer with absolutely no further cleverness required of the statistician that it's done once and for all all the cleverness went into designing the the idea and after that it's packaged and you do not have to do it anymore you don't have to think anymore and the bootstrap is the same way it's what I call an automatic theory an example of a non automatic theory that you may be familiar with is the theory of minimum variance unbiased estimation which you most of you have probably learned about and that is not an automatic theory for each new problem you have to cleverly try and figure out whether or not you can find a umv you estimate in that case or not if if that wasn't if the minimum variance unbiased theory was automatic I guarantee you it would have replaced maximum likelihood estimation as the dominant practical estimation technique it is not automatic automatic theories are very important for me when I'm working in a day to day mode trying to get somebody an answer I don't want to have to be clever every time a new problem comes in it's it's it's too much of a strain on my aging brain and I I really want to give them an answer in an automatic quick way that I can be fairly certain is close to optimal and that's just what both of these theories do talk about the simplest possible case first and this really doesn't work very well simple as possible cases the sample mean it's it's one that will we'll use to motivate the general theory so it we might have decided to answer the question about remember the fifteen observed lifetimes and question one was oh we were supposed to have a parameter of interest theta and we wanted to estimate it in question one was how should we estimated and let's just say that we've answered question one with the traditional answer molestin will estimate the expected lifetime with the sample expectation that is the sample mean and X bar that is just the sum of the N observations I cannot stop using n and is 15 here anyway there it is summation X I from one day and divided by N and if you look on that separate sheet of paper it says someplace that the mean is point 804 it's even in a little box does anybody not see where it says mean equals point 804 please tell me if you did not see my feelings will not be hurt I'll point it out to people okay and what's what's special okay so how active so now let's get on to question two how accurate is the sample mean for estimating the true mean we know perfectly well that the true mean of the population is not exactly point 804 well there's a simple formula for the standard error that everybody learned in their first statistics class and the simple formula expresses the standard error of X bar I've written that there a sigma of F and X bar with X bar in quotes let me say why I've written it that way Sigma stands for standard error and we'll all today f is the unknown sampling distribution from which we are drawing the exes and I'll say a little more about that in a second n is the sample size we know that 15 X bar in quotes is the way of forming the statistic I put it in quotes because it's not the number that we got point 804 it's the way we took a set of numbers and condensed them into one number and that we know also so we don't need to put in an X bar in our notation because we know an N X and quotes X bar we know that we know that we're using sample size n and we know that we're using the sample mean so I just call that Sigma of F there and the so Sigma of F is the standard error when sampling 15 times standard error of the mean when sampling 15 times from unknown population F and forming the sample mean and what we happen to know is that the that standard error has a nice formula and the nice formula is written the next line there it's the second central moment of F divided by n to the one-half power and this is really it really is a simple formula and everybody learned it a long time ago but the fact that it's a simple formula should not conceal from you that it's a very wonderful formula it's it's it's actually an integration over n dimensional space done very cleverly and when people realize this this was an important step forward in early statistics the second central moment incidentally if I was let me write it the second central moment is the expected value of the square minus the square of the expected value it's the second moment around the mean just for those who want to be reminded what's a random sample a random sample well we have a true distribution F and I've gone a little picture there of F yielding x1 through xn by random sampling by which I always a picture the population F being written out and lots of slips of paper a matter of fact an infinite number of pieces of paper because I like to think of infinite populations so we don't have to worry about finite population difficulties put in a hat and very well mixed and n draws in this case 15 draws made from from it it's what you might call in an iid sample and independent identically distributed sample a random sample the X's are independent they all have the same distribution and that distribution is F yes we feel yes in this case SiC Sigma is not the question was the question was is is does doesn't Sigma apply to theta there are theta hat is it yes yes thank you for ask thank you yeah thank you for am it for making that point I always use the term standard error for the standard deviation of a statistic of interest to make the distinction between that and say the standard deviation of any one of the individual x i's so yes Sigma is the standard deviation of what I call theta hat at first not the standard deviation of one of the X's that does that answer the point thank you for asking that because that that's a typical point that that the difficulty with giving talks on things you know or have given too many talks on is that you always forget the important points and that's one of them instead this the stand it's the standard error of the estimator or of the estimate of interest not of one of the individual components and that's what we're trying to estimate because that that's the thing that gives us an answer or at least a partial answer to question two thank you appreciate that okay how can we have the next slide please so we seem to have a nice formula for the standard error the one trouble with the formulas that involve something unknown it involves the second central moment of f so we don't seem to be any further ahead than we were before because we don't know that however it's quite easy to estimate the second central moment of F and that's what's usually done and in your elementary statistics class long long ago longer go for some than for others the you probably use the unbiased estimate of Sigma of the second central moment YouTube bar which is the summation of X I minus X bar squared divided by n minus 1 and I better bet they made a lot of fuss about dividing by n minus 1 instead of dividing by n because that's the thing that makes the estimate unbiased and then you just plug that into the formula for MU 2 for 4 Sigma of F for the unknown standard error and you get the usual estimate of of you get the usual estimated standard error for the mean which is the second central moment of the the estimated second central moment of the sample mu 2 bar divided by n to the 1/2 power and if I were to write that out it would come out summation X I minus X bar squared divided by n minus 1 and then divided again by n to the one half power and that's a formula I bet every one of you is used many times there's an easier way to get to to substitute into the formula for Sigma of F and let me describe the easier way but the easier way or the one that would have occurred I think to most of us first if we were faced with this problem say a couple hundred years ago is why not instead of s we don't know F but we happen to know F hat what's F hat F hat is the empirical distribution of the data the distribution that puts mass 1 over N on each of the endpoints and that's if you just draw the points on the line and make equal size dots at each of the points in a sense you're you're thinking of f hat in our in our particular case you put mass 115th at 0.143 mass 115th at point 1 8 2 dot dot dot mass by mass I mean probability one fifteenth at two point O seven six and that would be the empirical distribution we could why not here's an idea we have a formula Sigma of F for the standard error of the of the sample mean no matter what F is we don't know what F is why not substitute F 4 F hat for F that's a simple idea and I've done that you get almost the same answers above what what's the difference there's one difference between the between this answer and the answer above the yes the the the divisor is an instead of n minus 1 you get the maximum likelihood estimate instead of the unbiased estimate so if you now plug that in Sigma of F hat is almost the same as that's what I'll call Sigma hat and that's almost the same as what I call Sigma bar above the usual estimate except in the denominator there'll be instead of n minus 1 times n there'll be n times n and that's that's not a big difference of course it's not of practical there's no theoretical reason for preferring one to the other and as a matter of fact you can make arguments both ways you can make arguments that's good two divided by n plus one for example a decision theoretic kind of argument let me point out that the the second I I said it's easier the reason I said it's easier is that it's conceptually easier what in the first the first step was a very simple decision theoretic kind of argument that is the first way we did it where we use Mew to bar we actually thought about what was going into the formula and try to get a good estimate of something in the formula well that's that's not automatic if we have to think each time about what's a good estimate the problems are going to get a lot more complicated that's that's a clever and today clever will be a word with negative connotations that's a clever way to do the problem just substituting F hat for F has nothing special to do with this problem and in that sense it's conceptually easier so on to the next slide instantly that Sigma hat is the bootstrap estimate of standard error for the mean big deal huh why don't we just always substitute F hat for F anytime we have a problem and in Sigma of F well the fly in the ointment is that for more complicated statistics but there is no formula for Sigma of F and if if you think back to that course that I keep referring to you'll remember that somehow there weren't any other formulas like that like the one for the mean there wasn't you didn't get a formula for the standard error of the median the sample median and the reason is there is no formula there's approximate formulas and if you take more advanced courses you start learning more and more approximations formulas for standard errors for more and more complicated statistics but basically there is no simple expression like the one we had before so for example here's here's some examples we'll follow through the sample median for the 15 for the 15 life times the sample median was 0.6 1 1 and that's on your handout also it says median equals 0.6 one one it's in the little box up near the top the 10% trim mean was point 7 3 4 10 percent from mean means you take in this case one and a half observations off of each end and I won't say what taking a half of an observation off but said obvious then you take the the mean of the of the middle 12 in this case the remaining 80 percent of the data that was point seven three four we would like to know the standard errors of those estimates also however there is no simple formula for the standard error there's no formula for Sigma F and quotes theta hat where theta hat is the statistics say like the sample median again I'll just always call that Sigma of F so we can't substitute F hat for F the bootstrap what the bootstrap is is a numerical algorithm for always finding the numerical value of Sigma hat the bootstrap estimate which is Sigma of f hat it's just the formula if we had the formula Sigma of F with F hat substituted for F but you don't have to know the formula to do it it doesn't tell you it doesn't give you any analytic results what it gives you is the number that you would get and that's what the bootstrap is it's a good time to interrupt me stop me here if I've why is it called the bootstrap that's a good question and as a matter of fact originally I wasn't going to call it the bootstrap I was going to call it the combination distribution for reasons I'll tell you and I was really glad I didn't because it's called the bootstrap because in the original paper I was writing about the jackknife and I wanted something that didn't sound any more technical than the jackknife so I blame Tookie for the name the there's a serious reason why it's called the bootstrap and that is it as you'll see it involves reusing the data in a certain way that seems to be pulling yourself up by your own bootstraps that's an old story from Germany I think of Baron Munchausen pulling himself out of the mud by lifting himself up by his own bootstraps which of course is impossible and in fact the for a long time a bootstrap machine was the definition of an impossible physical device but in it's gotten I hope more positive connotations here here it's a device that that is indeed a self-help kind of thing and you'll see that we reuse the data okay to explain how the bootstrap algorithm works and now it'll become clear why why we call it the bootstrap I have to define what I mean by a bootstrap sample and a bootstrap sample will always you'll know we're dealing with bootstrap stuff because there's always going to be a star up on the top and the star which didn't come out awfully well and run with a fat pen indicates that actually this is not going to be real data it's going to be data generated from a computer or a random number device and that's where the randomness is going to come from a bootstrap sample y star equals x1 star through xn star is a random sample of size n drawn from F hat and I crossed out with replacement when Rob pointed out to me that random sample meant with replacement that is a random sample in my opinion in my definition in our definition here today is we assume the sample is infinite so even though F F hat is a distribution that only has n points support points we still think of it as an infinite population and we draw n times from that it's equivalent to say that we've drawn a sample of size n with replacement from the original set of numbers x1 through xn so for example we might get if we you you you can think of writing the numbers x1 through xn like the 15 lifetimes you can think of writing them each of them a million times and putting them all on a hat so you have 15 million slips of paper and the hat mixing up real well and drawing a sample of size 15 from that that would be a bootstrap sample yes yes yes now the question is is the why do we choose the bootstrap sample size to be the same as the original sample size and that that's a question that is asked of me very often and the the answer is is simple remember it's Sigma of F and theta hat that we're trying to estimate and I've suppressed N and theta hat from the notation but we used the same statistical form theta hat we also used the same and if we try and use a different size biases we'll be introduced that have to be corrected later that is we will not estimate Sigma of F hat sometimes it's conceivable that that might be a useful technique and then to recorrect later on but in fact for I've never found any advantage to using any sample size except in so the answer to the question is why is n in the bootstrap the same as anders before is that we're trying to mimic the real situation as closely as possible in order to get sigma of f hat we want to get sigma of f hat is really sigma of f hat and theta hat thank you another good list ok next slide please so that's what a bootstrap sample is and here's how the bootstrap algorithm works step it's a three step algorithm and I'm never any good I can't use the new algorithmic notation there's neat ways to write these things and Rob's real good at them but I can't ever write them right so I just write my algorithms out in line and I don't have it written right but anyway I think it'll be clear independently draw capital B bootstrap samples capital B is going to be a number like a hundred or a thousand and we'll talk about how big it has to be let's say it's a hundred right now okay yes question yes let me say that this the question is are these subsets of the original samples we're drawing each each bootstrap sample let's look at y star one y star one is a vector it has an objects in it X 1 star X 2 star xn star these n objects were gotten by drawing with replacement from f hat which is to say that we put the numbers we took the set of original values x1 through xn and drew and x with replacement remember I'm not saying without replacement I'm saying with replacement and so for example the first number might have shown up three times in the bootstrap sample the second number might not have shown up at all as the third one might have shown up once etc okay having done that once that gives us Y star one now we start the whole thing again we get Y star two finally we get Y star capital B so we'll have drawn B times n little n times so if B is 100 and n is 15 there have been a total of 1,500 draws to construct those 100 samples and these here we are answering your question that's why it's called the bootstrap we're using the data again yes the question is is is there a limit to how big B can be as a matter of fact there's there's a theoretical limit of how many different bootstrap samples there are with N equals 15 there's actually 76 million about seventy six point two million different bootstrap samples possible they're all the combinations of 15 things taken from a set of 15 but we don't want to do 76 right to million even even with a modern computer that's going to take a little too long so that's why we're doing random sampling yeah you can start getting the same some of the wise Darby's might be the same that won't infect things but it in fact I have never gotten the same bootstrap sample I always work with samples of size 10 or something like that or in my examples and I've never ever gotten a duplicate bootstrap sample that I know now I might have because I haven't looked at the data the the output carefully enough sometimes it doesn't affect the theory though that we might get the same one these are good questions today okay for each for each bootstrap sample we can re-evaluate the statistic of interest for example let's let's concentrate on the sample median we could re-evaluate the sample median again and again so for y star one we could get the sample median for y star two we could get the sample median and we'll call those the bootstrap medians or the bootstrap replications of the statistic in general so now we have say a hundred bootstrapped replications of our original statistic and then we simply calculate the empirical standard deviation of those numbers of those 100 numbers or capital B numbers that is we take the average of the theta hat stars and there it is I've called it theta hat star dot just to indicate that we took the average of the B numbers and then we formed the the root mean square sum affairs and I've Drew and I've divided by B minus 1 instead of B just to B consonant with the usual literature it wouldn't make any difference that's a number I call Sigma hat sub B it's not quite the same as the number I called Sigma hat which is Sigma of f hat what would make it the same as Sigma of f hat if I if I let B be infinity if I let capital B be infinity or actually in if I was clever if I let it be seventy six point two million and took them all carefully you not by random sampling then I could actually get Sigma hat infinity of course that which would be Sigma of F hat we won't do that because that takes too long instead we're going to be stuck with something like Sigma hat sub 100 which means the bootstrap estimate based on 100 bootstrap replications and I'll turn out that's fine we'll do some calculations to show that we have the next page placement I've simplified the algorithm looked messy on the previous page in fact it's really quite simple and I've written it down a little more simply here F hat gives y star B by random sampling Y y star B gives theta hat star B by about reevaluating the statistic and then the reevaluated values give Sigma hat of B by the usual formation of a standard sample standard error and that's all there is to it let's let's look back at oh the next slide okay as as capital B goes to infinity then Sigma hat of B goes to Sigma hat the ideal bootstrap estimate of Sigma which is Sigma of F hat which is Sigma of f hat and theta hat and there's dangers in boiling down the notation as much as I did to say Sigma hat is it there's a danger of losing track of what you estimated and what you knew but in fact all we did is we knew and we knew that the forms theta hat all we did was substitute F for a F hat for F and that's Sigma hat except we couldn't really get that so we approximated that by Montecarlo and that's the actual bootstrap estimate that you usually have to use Sigma hat sub B at least in a nonparametric situation if you look at if you'll go back and look at the the hand out that said bootstrap methods Efrain and tip shirani I've actually gone through this for the median and let me let me put this back over here for a minute the near the bottom of that page it's iíve shown you the first ten bootstrap samples for well they're the first ten bootstrap samples I forgot to say this it doesn't matter what statistic you're going to evaluate the bootstrap samples are always the same it's what statistic you evaluate over them that's different that's handy because that means you can write a computer algorithm that's very general for getting getting the answers and let's look at the second bootstrap sample do you see down at the bottom I've outlined where it says the second bootstrap sample does anybody not see where I said where it says the second bootstrap sample because it's fun I've ordered them I've reordered the numbers of course they didn't come out in proper order when I did random sampling but i reorganize x appears once in the sample point one four three appeared once in the second bootstrap sample point one eight two did not appear well that happens matter of fact it happens about 32% of the time that a number will not appear you can figure out what the probability is later if you're in the mood to do that neither the point 2 4 5 6 or 0.26 though or point oh yeah point 2 7 oh did appear that's the fifth one point six point four three seven appeared 0.55 oh nine did not appear six one one did appear oh excuse me yeah point six point one did appear point seven one two appeared four times not not an unlikely event at all you can figure that out etc and there of course we got sample size 15 that's a bootstrap sample and I've explicitly written out the first ten of them here just because it's it's sort of fun to look at them and see what they see what they look like and also in the case of the sample median it's very easy to evaluate the statistic once you have the ordered data because the statistic is always just the eighth-largest that that's unusual for the meet the median is on you statistic in that it's very easy to evaluate from the sample so so in this case theta hat star was 0.71 - it wasn't point 6 1 1 which is the median for the original data that's why I didn't look at the first bootstrap sample it happened to come out to be the same as the original data point so that would've been much fun the median is unusual in that the bootstrap median also always has to be one of the original data values it could be the smallest data value right could we could have gotten point one four three what would it take for in order to get point one four three what would have had to happen yes the we would have had to get eight or more observations on point one four three for that to be the bootstrap media and that doesn't happen very often is you can figure out never happened in those I actually did a hundred of them and it never happened okay we did this a hundred times and up at the top again it says bootstrap the bootstrap Sigma hat Sigma had a hundred equal to point two two nine I did it for the mean also and it came out point one five six that's up near the top of the handout there's the there's the bootstrap Sigma hat is point one five six for the median it came out point two two nine Oh in case you think taking medians always reduce the standard errors for distributions that look like they have a long tail it certainly doesn't for this this case in fact you've increased the standard error quite a bit Sigma and both of those are based on a hundred for the for the mean if you'll remember we happen to know Sigma hat infinity it's just what I call Sigma hat before because there's a close there's a simple formula that number is actually 0.155 here so you can see we're pretty close for the median we do not have the simple formula so we do not know what Sigma hat infinity is however in the middle I've traced what happened to the bootstrap median as we went along as B got bigger what happened to the Sigma hat for the median as B got bigger at at for vehicles 25 it was 0.23 for vehicles 50 was 0.224 B equals 100 it was 0.23 that is actually 0.229 I gave you more digits above yeah when we got up to a thousand was 0.25 and it's matter of fact has pretty well settled down by them so you can see that a hundred is not a bad number to use here question yes yes I do and it's coming up I the question was do I have do I actually know what F was in this case is it real data or is it simulated data its simulated data and the reason was that I wanted to be able to tell you what the correct answer was here but I didn't want to tell you too early because if I tell you too early as soon as I say what the distribution is that it was drawn from everybody else a well why didn't you use the right answer for that one and the answer and even if you don't say it you'll think it because it's impossible not we're trained to think that way they you know well if you know does that say the distribution is a one-sided querido why didn't you use the optimum answer there well we don't in real problems you don't know what distribution most times you do not know question suppose the original 15 observations were serially correlated that example will be the subject of what we're going to talk about this afternoon yes question the the question is I I said that as B goes to infinity you actually get to Sigma hat which is Sigma of f hat that's that's a that's a function of the strong law of large numbers and we could prove that easily how quickly it happens is a matter of error analysis and later on today I'll show you the error analysis to show you how quickly you get there in fact it always turns out because of the simple nature of the bootstrap you can say that in most cases you'll get there rather quickly it's its tenth Oh the other question is does it depend on n it it depends more on the form of this of the statistics statistics that tend to produce straggly outliers which are very sensitive statistics very non robust statistics it tends to take larger values of B though in fact if then you don't have to use the usual formula for calculating Sigma hat you can instead try and use a formula that's more robust to calculate Sigma hat B and get around that problem in fact a fairly safe general statement is B equals 100 is almost always enough for most standard error calculations regardless of an regardless of the form of the statistic if you're just a little cautious about how you how you estimate Sigma hat if you don't blindly plug into the formula summation theta hat star B minus theta hat star dot squared etc it really is a good idea to plot the histogram of the theta hat star values the bootstrap histogram and actually look at it and to see because if there's one of them that's far away you don't want to really include that in the in the formula for the variance near the middle way in the back read read
Online Copy: https://www.youtube.com/watch?v=4gjHv-Yry14
Metadata Source:YouTube
No holdings listed.
No related films.
Record added: 2026-06-01 13:26:50