Showing posts with label talos. Show all posts
Showing posts with label talos. Show all posts

Friday, December 11, 2009

So you want a new Talos suite, eh?

This quarter Alice and I have focused on trimming the list of pending test suites and where several new ones (419776, 524089, 515540, 506772) have been turned on in production.

The process for getting a new suite in has becoming a lot clearer, so we gave a presentation at the recent all-hands to help the developer know what to do on their end and what RelEng can do for them once their test suite is ready for staging.

Here's what a developer needs to do:
  • Download and install Standalone Talos to test their suite in
  • Once they have established that the test works on at least one platform, write a patch against talos in cvs
  • File a bug against RelEng in the General component and provide the following information:
    • Contact person who will work with us on getting the test suite enabled
    • What the test does, what the expected output should be, long name, short description
    • Which branches and platforms you want the test run on
RelEng will create the buildbot patches that enable the tests, insert the tests into graph server, and work with the contact person while the tests are in staging to make sure the expected outcome is reached. Once the tests run as expected we can turn them on in production. Perfect world turnaround for this process is about a week and a half and involves a short Talos downtime. The rest of the time allotted to our presentation was spent discussing where we should be setting our sights for Talos improvements. This looks to involve two relatively large undertakings:
  1. While it recently underwent some much needed improvements, the graph server still needs to be faster, more stable, scalable, and able to handle our ever-growing data sets. The blocker here is that no one really owns graph server and it's hard to know who should.
  2. Talos is barely holding up under the current load of tests, hardware, and infrastructure. It also works in such a way that a lot of manual involvement is required to add new tests. It would be awesome for it to work more like unittests where once individual tests are checked in, they would go into production immediately. It would then be possible for a developer to not only write a unittest for any bug fix, but also a performance test to go along with it.

Now this brings up the problem of what performance we want to measure and how we want to approach performance metrics in the long run. Alice made a great point when she stated that folks who are not trained and accustomed to doing QA might be challenged by trying to generate tests that actually create a good metric for the performance they wish to be testing. It's entirely possible to have tests that seem interesting on the surface, when you drill down, don't provide any useful data for actually improving anything.

Do we want per-bug performance tests as we do with unittests? While it looks like this is a way to make a developer more accountable for their code, it's pretty obvious that this model wouldn't scale well at all with our current hardware and turnaround expectancy. Imagine as many individual pageload tests as there are mochitests...I suspect no one wants to see that.

Performance testing would be better and more useful if it was targeted at specific features or areas of the product where someone is actually tracking the improvement/regression ranges on them as they are developed. That's a key area of Talos - that a human is actually accessing the data, finding it useful, and making improvements on their feature/area as a result of this information.

While brainstorming with Aki on the potential of the graph server data, one idea really got me excited. Open up the data.

There's been a lot of hype lately about opening up data. In February of this year Tim Berners-Lee encouraged us to start thinking about open, linked data and how it could be the next round in how the Web helps us re-frame our world. In Canada the city of Vancouver opened up its data in the hopes of "improving liveability and governance" in the Metro area.


What if the Talos graph data was made available to the community and a challenge was created in the spirit of the marketing design challenges where we ask people to help us find new ways to view the data? I'd be really curious to see what kind of visualizations would come out of the larger community. RelEng doesn't have a very large community outside of employees, so this could be a great way to start working on creating one.

Friday, October 23, 2009

Upcoming improvements to Talos documentation and test suite creation

This quarter I'm going to be joining Alice in trying to improve the system for adding new suites to Talos.  The current system involves a lot of hackery on our side and slows down the ability for us to get Talos suites up and running as quickly as might be desired.

So with John's help to create a prioritized list of suite requests, we will be doing a lot of communicating with developers in the coming months to get them up and to improve the process and documentation at the same time.  Currently there are 10 new suite requests waiting that are known and there may be others.  

Part of the issue with adding new suites is that there is a lack of documentation and tools for developers.  Our new system will look more like this:

* A request is made for a new suite and a developer is attached to the request who will be the lead person for working with us to get the suite into production

* The dev will be able to use tools we provide (standalone talos, corral of staging-talos slaves) to do proof of concept on the suite so that it works and is ready to go up in staging when it's handed over to RelEng

* RelEng will enable the test suite in staging and verify that changes in staging work fine with the other existing jobs being run on the same machines. Once all is well, then rollout to production would happen

As we progress through the suite requests, this process should get easier for all parties and more streamlined.  We hope that by the time we reach suite #10 it will be much easier and faster for developers and RelEng to get the proposed new Talos suites into production.

I mentioned the developers will have tools provided by us. We need to do a bit of work to make these tools usable by developers and the first place to start is with our documentation of what Talos is and how it works.  Following this we will have discussed having boilerplate code for creating each of the two styles of tests startup or pageload.  Also, it might be beneficial to have a coral of Talos machines that can be loaned out to a dev for a limited time in order to test a suite during creation and debugging.  This coral could then be re-imaged and passed along to the next suite developer.

Here is the current documentation page.  Doesn't give you much to go on, right?

Well this is about to change.  Given my complete lack of Talos knowledge, I will be writing up what I learn about Talos as it's happening so that hopefully a more complete set of docs will exist for the Talos neophyte and folks who want to work with us to add new suites will benefit from this as well.

Here's the current list of the docs to be created based on what we think you might want to know:

* How Talos works and an overview of the development from past to present

* What preferences Talos runs with

* A description of each test suite, what each runs

* What the numbers mean

These are the things I don't know - is there anything you don't see listed here that you want to know more about?  Feel free to make suggestions in the comments.

Tuesday, April 1, 2008

Build & Release - Learning about Talos

Armen and I are in California attending Build and Release team meetings this week. Over the next two days we'll be introduced to the many facets of the Build and Release workflow.

Todays first session was about Talos with Alice.

Here is the diagram of Talos (copied from the diagram Alice drew - yes Talos is a robot):


Alice walked us through how Talos gets its information from browser builds by having buildbot read for new builds from quickparse which is a text file. Buildbot has a script that knows what to look for in order to find new builds and there is a 5 minute delay before Talos is deployed because the information can get into quickparse before the build is finished and so therefore does not technically exist.

Currently there are 30 Production and 20 Stage Talos machines running, this past December there was only 1 Production machine and the stage machines.

This huge increase of Talos machines has led to an insane amount of data being gathered and a database which is in serious need of some help.

After Alice's presentation we all tried a standalone Talos so we could see the tests at work.

Anyone can try them, just follow these directions. If you are using a recent nightly you might need to add security.fileuri.strict_origin_policy : false to your sample.config file preferences because of new security features. Also, you can comment out the tjss tests because those are kind of boring - the fun test is the svg since you'll see a lot of graphics tests running on your browser. This standalone runs on a new profile so it's okay if you already have Firefox running when you run this script.

So the information that's generated is good for recognizing regression like in this bug, where if you look at the graph you can see how the build was chugging along, something got checked in that affected performance and then it was backed out and the performance went back to normal.

Pic here, see bug #425941 for more info:


More information on Talos Machines.