Sunday, July 05, 2009

Regression of 5D solubility space and distributed automation

I recently reported on the plotting of a solubility surface in 3D. Marshall Moritz has now extended his measurements of the solubility of 4-nitrobenzaldehyde in 2 more mixed solvent systems (ONSC-EXP114), giving us 4 solvents and temperature. The results are stored in the SolSumMix spreadsheet.

Andrew Lang has performed a quadratic regression analysis of this space and we have pretty good agreement with the experimental data points (see the "predicted solubility" column in the above SolSumMix spreadsheet).

Although we can't easily represent the entire 5D space intuitively, we can take 3D slices of the regression to assess the fit. For example, consider the plot of mol fraction % chloroform vs. acetonitrile keeping other solvents at zero concentration. What we observe is a nice saddle shape similar to the plot we did earlier with the original data points.
Now consider a slice of mol fraction % toluene vs THF keeping the other 2 solvents at zero. For temperatures above about 0 C we observe an expected rise in solubility with temperature. However going below 0 C the curve reverses and solubility is predicted to go up a bit. This is clearly not right and it simply means that we are missing key data points in that area. A quadratic fit will insert parabolic elements giving this inversion. It is very important to understand that these models will probably do fairly well for intrapolating within our experimental range (about -25C to 40C) but will not be very helpful for extrapolating beyond this region.
Since it is difficult to manually inspect every possible slice of this 5D space Andy has created a service that returns recommendations for the most needed points to be measured next to generate a better model. We already have a DoSol spreadsheet that instructs ONS Challenge students as to the most urgent next solubility measurement to make. This additional "bot" (not quite fully automated as of yet but will be soon) integrates nicely with the collection we already have.

We aim to show that such open distributed mechanisms to requests and execute measurements is a viable way to efficiently leverage crowdsourcing to automate parts of the scientific method. If it can be applied to solubility it can be applied to other problems.

Labels: , , ,

Tuesday, March 03, 2009

Semi-automated measurement of solubility using NMR

Over the past few days Andrew Lang and I have been discussing ways of streamlining the measurement of non-aqueous solubilities using NMR. Inspired by David Strumfels' VBA code on Excel to automatically measure kinetics, Andy found a way to directly extract the integration values from the H NMR spectra hosted on our server in JCAMP-DX format.

We have set up a Google Spreadsheet (see ONSC-EXP062B for an example) that automatically calculates solubility based on information that the researcher provides. What is required:
  • A link to the NMR spectrum (the HTML file linking to the JCAMP-DX file)
  • Density and molecular weight of the solute
  • Density and molecular weight of the solvent
  • A range in the solute to integrate with the number of corresponding Hs
  • A range in the solvent to integrate with the number of corresponding Hs
The spreadsheet calculates the molar ratio of the solute to solvent then the molarity by making use of the assumption that the volumes of the two components are additive. The volume of solids is typically not available experimentally but ChemSpider gives a reasonable prediction. We have been using this assumption to convert published solubility values from g solute/ 100 g solvent to molar. In order to prevent taxing the server, once the measurements are computed they are stored in a database so that the spectrum and calculations don't have to be performed again.

The beauty of this approach is that there are no volume measurements. A saturated solution is made then, generally diluted in a deuterated solvent. When using an internal standard the volume of the saturated solution and the volume (or weight) of the standard must be known exactly. It is often difficult to micropipette some solvents and there is always the possibility of making an error in the handling of the micropipette. In general the fewer variables there are the more likely the results will be reproducible.

This is method can save a lot of time but it is not as automated as it could be. The densities must be looked up manually, although the molecular weight is automatically calculated from the common name using a web service by Rajarshi Guha run directly from within Google Spreadsheets. It also requires students to define solvent and solute ranges manually. All of the input cells are colored green, the output red and the intermediate calculations are in yellow.

However, once a range for a solute or solvent (and corresponding number of hydrogens) has been determined it can be used as a handy default and we will be collecting these and storing them in this sheet.

Does this mean that students don't need to think anymore?

Used properly, this system should actually elevate the level of thinking, in much the same way that the calculator did not remove the need for thought in data analysis. It just removed a lot of the tedium of manually calculating square roots and all of the associated sources of error in manual calculations of that type.

Students should use this tool to handle more measurements - faster - and think about their results in aggregate form. The ability to detect systematic errors becomes an essential skill to be developed. Also students need to spot problematic results quickly, for example where solvent and solute peaks overlap - or where there are baseline anomalies.

Of course even these last issues of quality can control can probably be automated to a large extent and we will report on this as we go. For example, it is conceivable that the NMR of a solute and solvent can be predicted or looked up automatically (on ChemSpider for example) and probable peak overlaps could be flagged. Software could also probably detect a mislabeling error.

At this point, we are getting closer to scientific progress by machine-to-machine communication on the free open read/write web. All we would need are a few groups around the world who see the value in endeavors such as this and donate a part of their NMR autosampling time. I am sure we could come up with simple ways of automatically converting files on their local computers to JCAMP-DX format and automatically upload them. We also need people to make up saturated solutions - but with the decoupling of tasks in this new workflow - these don't necessarily have to be the same people who process the NMR spectra.


Labels: , , ,

Monday, June 23, 2008

Laboratory Automation Talk at Drexel

Tom Osborne from Mettler-Toledo will give a talk on laboratory automation at 10:30 a.m. on Tuesday July 1, 2008 in Disque 109 (32nd and Chestnut streets).

The presentation will focus on chemistry applications of the MiniBlock, MiniBlock XT, and MiniMapper systems, including some results from the Bradley lab on the optimization of the Ugi reaction. A laboratory demonstration will follow in Disque 511.

RSVP Jean-Claude.Bradley@drexel.edu to attend.

MiniBlock® is a flexible, easy to use tool designed for parallel synthesis. MiniBlock® is the only compact parallel synthesizer that allows synthesis via Solid Phase or Solution Phase to be carried out on the same platform. Originally designed by medicinal and combinatorial chemists at Bristol-Myers Squibb Company, the MiniBlock® has been further developed to address a wide range of chemistry methodologies.


Labels: , , , ,

Sunday, June 15, 2008

Ugi Reaction on MiniMapper - Trial 1 Preliminary Results

After receiving the MiniMapper two weeks ago we now have some numbers to look at from Trial 1.

The idea here was to repeat a Ugi reaction known to work in methanol at 0.5M concentration (EXP099) and explore the experimental space varying solvent, concentration and an excess of some of the reagents.

All 48 reactions gave precipitate (EXP189). We still have to optimize the lighting and camera angle to see this clearly for all tubes:

In reviewing the results I noticed an inconsistency between the experimental design and the information in the GoogleDoc used to plan the details of the experiment. But because the MiniMapper stores the log of its actions in an XML log file, Khalid was able to make corrections to the worksheet.

Based on the weight of product and the limiting reagent, a yield for each reaction was calculated (column AC).

The best 4 yields (48-61%) were found for 4:1 acetonitrile/methanol solutions at the lowest concentration (0.1 M). It is counterintuitive that a multi-component reaction should work better at lower concentration but this is an effect that I noted previously.

Ironically, the absolute lowest yield (1.4%) was found for the 0.5M methanol reaction that was supposed to be the positive control!

This may be due to a malfunction in the delivery of one of the reagents. However, the other reactions carried out at this concentration in methanol were all low yielding (10-15%), compared to when it was run manually in EXP099 (37%).

I must stress that these results are very preliminary since we have not yet shown that the isolated materials are all pure Ugi products. As the NMRs come in I'll update the status.

But I think that this was a good first trial to demonstrate what kind of questions can be asked with the power of automation and the importance of reporting automated logs in Open Notebook Science.

Labels: , , , , , ,

Saturday, May 17, 2008

UsefulChem Automation Trial with Mettler Toledo

Kevin Owens and I have been looking into equipment to automate some UsefulChem experiments.

This is something that I feel strongly will become important in Open Science applications, especially as it relates to Open Notebook Science. I think it is one of the paths of least resistance for the automation of the scientific process.

Industry is automated to the gills but it will probably be easier to convince academic practitioners of Open Science to automate their procedures rather than to get industry to open their data. Can you imagine a company allowing crowds to design and analyze experiments run on their machines? That is what we've been proposing and it would be difficult to reconcile that with a business model based on IP protection.

In that NSF proposal, we planned to use ChemSpeed's technology. Kevin and I recently visited ChemSpeed at their Princeton location and we were impressed with the capabilities of their reactors. We're in the process of planning a trial run of the Ugi reaction on their system and we'll post on the progress of that on this UC wiki page. The idea is to couple a digital camera within the robot's workflow to be able to generate results comparable to those manually generated by my students.

ChemSpeed's systems are quite powerful but also expensive (200-400K). In order to take advantage of more funding opportunities, we've also been looking at Mettler-Toledo's MiniMapper/MiniBlock solutions. We're planning this out openly on this wiki page - any feedback is welcome.

I've had a good discussion with Frank Schoenen at the University of Kansas, where they run both systems as part of servicing the NIH Roadmap Program. Based on his feedback I think this trial run should be successful.

Labels: , ,

Thursday, January 03, 2008

Modularizing Results and Analysis in Chemistry

Chemical research has traditionally been organized in either experiment-centric or molecule-centric models.

This makes sense from the chemist's standpoint.

When we think about doing chemistry, we conceptualize experiments as the fundamental unit of progress. This is reflected in the laboratory notebook, where each page is an experiment, with an objective, a procedure, the results, their analysis and a final conclusion optimally directly answering the stated objective.

When we think about searching for chemistry, we generally imagine molecules and transformations. This is reflected in the search engines that are available to chemists, with most allowing at least the drawing or representation of a single molecule or class of molecules (via substructure searching).

But these are not the only perspectives possible.

What would chemistry look like from a results-centric view?

Lets see with a specific example. Take EXP150, where we are trying to synthesize a Ugi product as a potential anti-malarial agent and identify Ugi products that crystallize from their reaction mixture.

If we extract the information contained here based on individual results, something very interesting happens. By using some standard representation for actions we can come up with something that looks like it should be machine readable without much difficulty:
  • ADD container (type=one dram screwcap vial)
  • ADD methanol (InChIKey=OKKJLVBELUTLKV-UHFFFAOYAX, volume=1 ml)
  • WAIT (time=15 min)
  • ADD benzylamine (InChIKey=WGQKYBSKWIADBV-UHFFFAOYAL, volume=54.6 ul)
  • VORTEX (time=15 s)
  • WAIT (time=4 min)
  • ADD phenanthrene-9-carboxaldehyde (InChIKey=QECIGCMPORCORE-UHFFFAOYAE, mass=103.1 mg)
  • VORTEX (time=4 min)
  • WAIT (time=22 min)
  • ADD crotonic acid (InChIKey=LDHQCZJRKDOVOX-JSWHHWTPCJ, mass=43.0 mg)
  • VORTEX (time=30 s)
  • WAIT (time=14 min)
  • ADD tert-butyl isocyanide (InChIKey=FAGLEPBREOXSAC-UHFFFAOYAL, volume=56.5 ul)
  • VORTEX (time=5.5 min)
  • TAKE PICTURE



It turns out that for this CombiUgi project very few commands are required to describe all possible actions:
  • ADD
  • WAIT
  • VORTEX
  • CENTRIFUGE
  • DECANT
  • TAKE PICTURE
  • TAKE NMR
By focusing on each result independently, it no longer matters if the objective of the experiment was reached or if the experiment was aborted at a later point.

Also, if we recorded chemistry this way we could do searches that are currently not possible:
  • What happens (pictures, NMRs) when an amine and an aromatic aldehyde are mixed in an alcoholic solvent for more than 3 hours with at least 15 s vortexing after the addition of both reagents?
  • What happens (picture, NMRs) when an isonitrile, amine, aldehyde and carboxylic acid are mixed in that specific order, with at least 2 vortexing steps of any duration?
I am not sure if we can get to that level of query control, but ChemSpider will investigate representing our results in a database in this way to see how far we can get.

Note that we can't represent everything using this approach. For example observations made in the experiment log don't show up here, as well as anything unexpected. Therefore, at least as long as we have human beings recording experiments, we're going to continue to use the wiki as the official lab notebook of my group. But hopefully I've shown how we can translate from freeform to structured format fairly easily.

Now one reason I think that this is a good time to generate results-centric databases is the inevitable rise of automation. It turns out that it is difficult for humans to record an experiment log accurately. (Take a look at the lab notebooks in a typical organic chemistry lab - can you really reproduce all those experiments without talking to the researcher?)

But machines are good at recording dates and times of actions and all the tedious details of executing a protocol. This is something that we would like to address in the automation component of our next proposal.

Does that mean that machines will replace chemists in the near future? Not any more than calculators have replaced mathematicians. I think that automating result production will leave more time for analysis, which is really the test of a true chemist (as opposed to a technician).

Here is an example of an analysis module making a simple point, useful to the chemistry community, and linking back to result modules that ultimately link back to the original experiment in the online laboratory notebook:
Context: obtaining precipitates in the CombiUgi project

Ugi reactions in methanol where the solution is supersaturated with Ugi product may give false negatives for precipitation. For example, a Ugi product rapidly crystallized at the 17th hour (RESULT0003) after addition of all reagents, while appearing as a clear solution at the 15th hour (RESULT0002). It is therefore recommended that the vials be submitted to vortexing (15 s) prior to taking a picture.
We'll be recording these analysis and result modules on UsefulChem wiki pages:
We'll be using InChIKeys for compact unambiguous identification of molecules (and convenient indexing in Google) and the terms in this post for action options. Anyone is free to automatically incorporate these in a database, as long as attribution is provided. (If anyone knows of any accepted XML for experimental actions let me know and we'll adopt that.)

I think this takes us a step closer from freeform Open Notebook Science to the chemical semantic web, something that both Cameron Neylon and I have been discussing for a while now.

Labels: , , ,

Creative Commons Attribution Share-Alike 2.5 License