hi folks in the previous video clip we created some artificial banking information as well as began to discover the data both on system as well as off platform to take a look at just how depictive it was to the initial production information and also the final thought we drew was that you might tell the very same story with the Personal privacy synthetic data as you can inform with the original data in this video we” re mosting likely to have a look at a second usage case and also we” re going to take that exact same synthetic banking information as well as publish it to talk GPT code interpreter to do some explore exploration some evaluation and afterwards eventually utilize it to aid notify our next synchronization on platform so allow” s start so as you may be mindful in chat GPT plus there is a beta version of code interpreter which allows you to post files so as you can see right here I have my bank advertising artificial data saved I” m mosting likely to provide it a few prompts on the initial one is um a little of details regarding the data set itself the UCI web site and after that I” m going to ask it to tell me about the information create some basic charts and also visualizations and afterwards inform me the three essential variables affecting variable Y which once again is the climate the campaign was a success or not so allow” s kick that off and take a pair secs so allow me stop below so chat GPT is just completed doing its magic here um and also as you can see on screen um it has informed me provide me a profile of the information set all the columns the variables it” s produced some fundamental graphs and visualizations um as well as it” s likewise drawn out the 3 key variables impacting my classification objective um and I think this is simply stunning due to the fact that when we speak about data democratization and also it” s one point getting premium quality information in the hands of all information consumers however once you get that information in their hands you also have to make certain that they can understand the information that they can assess the data and inevitably use the data to ensure that interaction between performance like code interpreter and synthetic information which clearly allows you to publish the Personal privacy conserve artificial data in such a way that you wouldn” t make with your personal privacy delicate production data it” s simply beautiful as well as for information democratization so now that we have that basic evaluation done let reveal you how I” m going to utilize code interpreter to confirm my next synthesization on platform so I” m going to ask I wish to up sample the period of hire my information set since it” s informed me it ‘ s among the crucial variables as well as to do so I wish to comprehend what the typical size of those effective telephone calls is and as you can see on display it” s 540 secs so allow me switch back on system and also upload this information collection here which is Bank advertising and marketing customized So based upon the info told to me by conversation GPT what I did was create one more categorical column called period High what I” m going to do utilizing the most AI platform now is to up sample the number of circumstances where call durations mosted likely to 404 540 seconds or over to make sure that we can analyze exactly how that impacts our category objective why so what you do here on the most Air Platform is rather amazing you can rebalance a column as well as I” ve picked certainly period high as well as in the data set that is indicated by true iefa telephone call is 540 seconds or longer as well as it is marked as real so I” m going to up example that to 50 and after that I ‘ m mosting likely to select my destination and also I” m mosting likely to introduce the task and also simply like the previous video clip allow me go right into a previously run Bank Market information set and also I” ll experience to the QA record so essential distinction right here we have our design QA report which demonstrates how very closely the model has actually learned the initial data and also in the data QA record which if you increased or Diversified the data will show you how the artificial information varies from the original information and also as you can see right here in the duration High we currently have about half of instances of high period calls in the synthetic data compared to about 11 in the original information as well as if I close below I can see my crucial variable y i IFD project was a success or not you can see exactly how having more High duration calls has impacted in a favorable method the variety of successes I have from my direct marketing campaign so I think that” s actually actually cool and also simply include on this point in the original video clip we demonstrated how you can make use of privacy say artificial information to inform the same tale as the initial production data and also in this 2nd video what we” ve done is shown the beautiful awesome interplay you can have with code interpreter and also and after that just how you can boost artificial information to begin telling a various tale from the original information and also that” s simply one of the reasons that we think synthetic data is far better than genuine data many thanks for seeing

