Monday, December 6, 2010

Basics of “Stratified Sampling”

What I call “stratified sampling” is in part a way of sampling, but more importantly it is a way of determining compliance with the calculated acceptance limits in a protocol. Stratified sampling can be applied to both swab sampling and rinse sampling (or possibly the combination of the two); however, for simplicity I will use examples from swab sampling as my illustration.
The “traditional” way swab sampling is done is to select representative sampling locations, focusing on worst-case (most difficult to clean or most likely to leave residue behind) sampling locations. Assuming the limits are calculated using a dose-based (maximum allowable carryover) calculation, then my requirement is that every sampled location must meet the calculated limit. For example, if my calculation results in a limit of 1.0 μg/cm2, then every sampled site must meet that criterion. This generally results in significant overkill in terms of total carryover, because the limiting factor in meeting my acceptance criteria is the sampling site with the highest residue value.
What is different in “stratified sampling”? One still samples the worst case locations, but determination of meeting the residue acceptance criterion is based on the total carryover. This, of course, assumes that the residues carried over from equipment surfaces are uniformly distributed in the next manufactured product. How can I calculate the total carryover? One way is to stratify the equipment surfaces. This can be done for an individual equipment item, or it can be done for an equipment train. For simplicity, the example I use will be for a single equipment item.
In stratifying equipment, I first define representative surface segments of the equipment, and then determine the surface area of that segment. For example, for a liquid blending vessel, I might identify these segments:
Dome 5,000 cm2
Sidewalls 31,900 cm2
Bottom 5,000 cm2
Baffles 5,000 cm2
Drain 100 cm2
Blades 1,000 cm2
Shaft 2,000 cm2
The total surface area is 50,000 cm2.
Then during my protocol, I want to make sure I sample at least one worst-case location in each and every segment. Note that in a traditional approach, I might not sample a location on the vessel bottom. However, under stratified sampling principles, I must select a worst-case location for sampling for the tank bottom.
Then during the protocol execution, I measure residues in all sampling locations. The next step is to multiply the highest value of any swab sample for any swabbed site within a segment by the surface area of that segment. This gives me a maximum possible actual carryover for that segment. Let’s assume the measurements in the second column below are the highest values obtained for a given segment. The maximum actual carryover for each segment is given in the last column.
Dome: 0.15 μg/cm2 X 5,000 cm2 = 750 μg
Sidewalls: 0.10 μg/cm2 X 31,900 cm2 = 3,190 μg
Bottom: 0.10 μg/cm2 X 5,000 cm2 = 500 μg
Baffles: 0.12 μg/cm2 X 5,000 cm2 = 600 μg
Drain: 2.60 μg/cm2 X 100 cm2 = 260 μg
Blades: 2.10 μg/cm2 X 1,000 cm2 = 2,100 μg
Shaft: 1.97 μg/cm2 X 2,000 cm2 = 3,940 μg
Based on these data, the maximum possible actual carryover to the next product would be the sum of all these values, or 11,340 μg. This is the value I compare to my total carryover limit (the total carryover limit is what I usually call L2). If the limit per surface area is 1.0 μg/cm2, and if the total surface area of the vessel is 50,000 cm2, then the total carryover allowed is simply the product of those two values, or 50,000 μg. Since my maximum possible actual carryover is 11,340 μg, the measured total residue is below my calculated total carryover limit, and I have met my residue acceptance criterion.
This seems like a lot of calculations; so what’s the point? Well the point is that three of the sample locations (drain, blades and shaft) would have failed under the traditional approach of having every sample meet the surface area limit of 1.0 μg/cm2. However, because those three locations represent such a low percentage of the total surface area, those higher results are possible without having the total carryover exceeded.
Now it is still a requirement that all surfaces be visually clean. Therefore, you can’t take this so far as to say that the measured residue for the drain was 15.2 μg/cm2, and I could still meet the calculated carryover. While it might be possible to meet the calculated total carryover limit with that high a result, it is likely that the drain would fail the visually clean criterion (since 4 μg/cm2 is a typical visual limit). So the use of stratified sampling can only take you so far.
Remember, that this method can only be used if the residues from equipment surfaces are uniformly dispersed through the next manufactured product. Furthermore, it is preferably only used proactively. That is, define the segment in advance, select the worst case location(s) in each segment, and write your protocol with this approach. It is also preferable that this approach be permitted in your cleaning validation master plan or high level policy.
The purpose of the Cleaning Memo is to present the basics of a stratified sampling approach for determining compliance with residue acceptance criterion in a protocol. This approach, while not commonly used, is based on good science and good logic.

More on “Stratified Sampling”

Last month we covered the example of applying stratified sampling to segments within a given piece of equipment. For this month, I will cover an example where it is used for a series of equipment items in a manufacturing equipment train (for finished drug product manufacture). Let’s assume for simplicity that there are three equipment items used for manufacture of the drug product. I’ll call them P, Q and R. I'm cleaning Product A, and Product B is the next manufactured product. Let’s assume the surface area of the three equipment items are as follows:
Equipment P: 100,000 cm2
Equipment Q: 85,000 cm2
Equipment R: 15,000 cm2
In a typical carryover calculation, I would calculate my limit using the total surface area of the equipment train to arrive at a surface area limit. Let’s say that result (based on the dose of the active in Product A, the dose of the drug product Product B, the total shared surface area, and the batch size of Product B) is a L3 value (limit per surface area) of 1.0 μg/cm2. I would then require in my protocol that each swab sample meet that limit of 1.0 μg/cm2. If I were doing a separate sampling rinse for Equipment Q, and if I used a sampling rinse volume of 50 L, I would set a rinse limit for that equipment item based on conventional rinse calculation:
L4 = (L3) (Surface area sampled) = (1.0 μg/cm2) (85,000 cm2) = 1.7 μg/mL
(Rinse volume) (50,000 mL)
I would then expect my rinse sample to meet that L4 limit. And those determinations are perfectly acceptable (and commonplace) ways of determining that I meet my acceptance criterion.
But, I can also use stratified sampling to determine compliance with my calculated L2 limit (the total carryover). Remember that there are conditions to utilizing stratified sampling in this way. The primary concern is that the residues carried over from equipment surfaces are uniformly distributed in the next manufactured product.
Continuing with the example I started with, the equipment train is stratified by the individual equipment items in that train. Then during my protocol, I want to make sure I sample all the worst-case locations in each equipment item. Note that this sampling is essentially the same sampling (locations and number of samples) as if I were not doing stratified sampling. I then measure residues in all samples. The next step is to multiply the highest value of any swab sample for any swabbed site within a given equipment item by the surface area of that equipment item. This gives me a maximum possible actual carryover for that equipment item. Let’s assume the measurements in the second column below are the highest values obtained for a given equipment item. The maximum actual carryover for each equipment item is given in the last column.
Equipment P: 0.15 μg/cm2 X 100,000 cm2 = 15,000 μg
Equipment Q: 0.20 μg/cm2 X 85,000 cm2 = 17,000 μg
Equipment R: 1.20 μg/cm2 X 15,000 cm2 = 18,000 μg
The total possible carryover under this scenario would be the sum of all equipment items, or 50,000 μg (50 mg, for those of you used to seeing smaller numbers). Going back to my original calculation of a total carryover limit, if my average L3 was 1.0 μg/cm2 and if the total surface area were 200,000 cm2, my total L2 limit would be 200,000 μg. Since my actual maximum residue value as determined by stratified sampling was 50,000 μg, my cleaning was effective because it was under the residue limit.
The value of this approach is seen in that with the example used, I would have failed the protocol (at least for Equipment R) with this data, with at least one location in Equipment R being above the calculated L3 limit of 1.0 μg/cm2. But using a stratified sampling approach, I still have a scientific and logical rationale for saying the carryover is less than the calculated total amount. Yes I would be happier to see all data points meeting the one L3 limit of 1.0 μg/cm2. But, I would also be happier if al
l my data points were below the limit of detection (LOD) by the best available analytical technique. But at least for most situations, that is not required.
Some of you may question the use of this technique, never having seen it before. Frankly, when I started as an independent consultant, I had not seen this “stratified sampling” approach. When I saw it, my response was something like “It’s not the typical method most pharmaceutical companies use, but it does have a scientific and logical basis for use.” Particularly with all the talk about wanting to be on a sounder scientific rationale, there should be no serious objection to this technique provided it is used correctly and in appropriate situations.
Realize that this technique offers an advantage to a manufacturer mainly when a smaller equipment item has a larger residue value. However, it also offers some advantages when the larger equipment item has the higher swab residue values. Reversing the second column values for P and Q in the example previously given results in the following carryover values:
Equipment P: 1.20 μg/cm2 X 100,000 cm2 = 120,000 μg
Equipment Q: 0.20 μg/cm2 X 85,000 cm2 = 17,000 μg
Equipment R: 0.15 μg/cm2 X 15,000 cm2 = 2,250 μg
In this case, the total carry is 139,250 μg, which is still below the acceptance limit of 200,000 μg.
However, it is possible to carry this approach only so far. Here is a third case:
Equipment P: 0.15 μg/cm2 X 100,000 cm2 = 15,000 μg
Equipment Q: 0.20 μg/cm2 X 85,000 cm2 = 17,000 μg
Equipment R: 5.20 μg/cm2 X 15,000 cm2 = 78,000 μg
I can total the carryover to get a value of 110,000 μg, and it looks like I will pass my total carryover acceptance limit of 200,000 μg. However, in this case it is likely that I will fail my visually clean criterion (with a swab value of 5.20 μg/cm2). In other words, the use of this technique is not completely elastic.
The purpose of the Cleaning Memo is to present additional examples of the use of a stratified sampling approach for determining compliance with residue acceptance criterion in a protocol. This approach, while not commonly used, is based on good science and good logic. This Cleaning Memo should be read in conjunction with last month’s Cleaning Memo.

Final Notes on “Stratified Sampling”

The first question should be an obvious one. If stratified sampling can be applied to segments within a given piece of equipment and it can also be applied to a series of equipment items in a manufacturing equipment train, is it possible to apply stratified sampling where I stratify both segments within one piece of equipment and a series of equipment items in a manufacturing train? And the answer is “Yes”, although it significantly complicates the pre-protocol work that must be done, as well as the calculations necessary in the protocol execution itself. However, there is nothing logically or scientifically invalid about applying stratified sampling principles in these cases (provided of course, the restrictions or limitations discussed in the two previous Cleaning Memos are considered).
The second question involves the use of rinse sampling and swab sampling for a stratified approach (the examples given in the previous months all involved swab sampling). I will consider three cases. Case I: Is it possible to use this approach where only rinse sampling is performed (on separate equipment items in a train)? Case II: Or, to use it in an equipment train where some of the equipment is sampled by swabbing and some equipment is sampled by rinsing? Case III: Or, to use it in an equipment train where a given equipment item is sampled by both rinsing and swabbing? The answer to all questions is “Yes”. The key is just to use the stratified sampling principles appropriately.
In Case I (everything is only rinsed), it is important that the rinse limit be established on carryover calculation principles. Then, the total actual carryover for a given equipment item can be determined by multiplying the concentration of the residue in the rinse solution by the volume of the rinse solution. The total actual carryover for the equipment train can be determined by adding up the actual carryovers for each equipment item. That total is then compared to the total carryover limit determined by the MAC calculation.
In Case II (some items in a train are only swabbed and some are only rinsed), it is important that both the rinse limit and the swab limit be established on carryover calculation principles. For equipment items rinsed, the total actual carryover for a given equipment item can be determined by multiplying the concentration of the residue in the rinse solution by the volume of the rinse solution. For equipment items swabbed, the total actual carryover for a given equipment item can be determined by using the calculations shown in the March Cleaning Memo. The total actual carryover for the equipment train can be determined by adding up the actual carryovers for each equipment item (whether it sampled by swabbing or rinsing). That total is then compared to the total carryover limit determined by the MAC calculation.
In Case III (some items in a train are sampled by both swabbing and rinsing), it is again important that both the rinse limit and the swab limit be established on carryover calculation principles. For each equipment item both swabbed and rinsed, the total actual carryover for that equipment item can be first determined by using the swabbing data to determine the maximum actual carryover for that item based on swabbing. Then for the same equipment item, the maximum actual carryover is calculated based on the rinse data (using the principles in Case I above). Assuming both the swab data and the rinse data each give results representing the total actual carryover, the larger of the two actual carryover results is then used for the total carryover for that equipment item. [Note that other things being equal, the results based on swabbing will ordinarily give a higher result than the data based on rinsing.] This is done for each equipment item, and the total actual carryover for the equipment train can be determined by adding up the actual carryovers for each equipment item (however that item is sampled). That total is then compared to the total carryover limit determined by the MAC calculation.
This sounds like a lot of work, and it does represent extra calculations. However, there is another alternative in the use of stratified sampling. That alternative is to use a staged approach to determine whether the acceptance criterion in a protocol is met. This staged approach involves a initial evaluation of every swab and rinse sample result, and comparing it to the acceptance limits for swabbing (perhaps based on a limit expressed as μg/cm2) and for rinse samples (perhaps based on a limit expressed as ppm or μg/ mL). This is how it is ordinarily in a cleaning validation protocol. If all those results are at or below the acceptance limit, then there is no point in going further in stratified sampling; the residue limit criterion is met. However, if one data point (or more) for a swab or rinse sample is above the limit, then the next step is to proceed with a stratified sampling approach to see whether the total carryover is acceptable. In this case, it is preferable not to say the initial evaluation “failed” the acceptance criterion. It is better to say something like “Stage 1 criteria were not met, and we will proceed to a Stage 2 evaluation to determine acceptability”. If the stratified sample approach demonstrates that the total carryover was acceptable, then the residue limit criterion was met.
In this staged approach, one cannot have “unacceptable” results from a typical (Stage 1) evaluation, and then decide that it might pass by a stratified sampling evaluation. This staged approach should be written into the protocol. Furthermore, if segments within an equipment item are to be stratified, those segments should be identified in advance (in part to prevent the temptation to “adjust” the segments or segment surface areas based on the data obtained, so that the end result is more likely to meet the limit based on stratified principles).
This staged approach should not be foreign to those in pharmaceutical manufacturing. It is something we use on a regular basis in the USP testing for conductivity in Purified Water and WFI systems.
The purpose of the Cleaning Memo is to address additional issues in the use of a stratified sampling approach for determining compliance with residue acceptance criterion in a protocol. This approach, while not commonly used, is based on good science and good logic. I should reiterate some conditions for utilizing stratified sampling. This method can only be used if the residues from equipment surfaces are uniformly dispersed through the next manufactured product. Furthermore, it is preferably only used proactively. That is, define the segments in advance, select the worst case location(s) in each segment, and write your protocol with this approach. It is also preferable that this approach be permitted in your cleaning validation master plan or high level policy. Furthermore, this Cleaning Memo should be read in conjunction with the March and April (2010) Cleaning Memos.

Acceptable Variability for Sampling Recovery Studies

Several months ago (January 2010), my favorite statistician, Lynn Torbeck*, published an article in Pharmaceutical Technology entitled “%RSD: Friend or Foe”. The article was basically about the misuse of statistics. In it Mr. Torbeck made the statement that applying a percent relative standard deviation (%RSD) criterion to percentage recovery values in recovery studies was not statistically valid because the values themselves were already percentages. Now, it is common practice in cleaning validation programs to include a criterion for %RSD for recovery percentages in sampling recovery studies (such as swab recovery studies). For example, companies might specify for swab recovery studies that the data collected for a given spiked level have a %RSD of ≤15%.
This got me thinking. Is Lynn right, or are most of the pharmaceutical companies (as well as yours truly, who has taught the use of a %RSD criterion for recovery studies) right? Furthermore, if we abandon the %RSD criterion, what criterion do we use to measure variability in a sampling recovery study? And, finally, does it really make that much difference? In situations like this, I often try to look at practical data to see the impact.
Here is my first example, which in comparison to a second case, illustrates that %RSD is not necessarily a good measure of variability in a sampling recovery study. For simplicity, I am only going to illustrate this with swab sampling, involving one operator who performs three replicates. In Case A below, the 100% recovery value is 2.03 µg/cm2. The values obtained, the average, the standard deviation and the %RSD is given in the Table A below. For simplicity, the data on recovery values will omit the units (µg/cm2).
Table A Data
  Data values % Recovery Values
Replicate 1 value 1.93 95
Replicate 2 value 1.96 97
Replicate 3 value 2.01 99
Average 1.97 97
Standard deviation 0.040 2.0
%RSD 2.1 2.1
Now, we have a second operator with the following data for the exact same sampling situation. Table B has the data for that operator in this situation.
Table B Data
  Data values % Recovery Values
Replicate 1 value 1.03 51
Replicate 2 value 1.06 52
Replicate 3 value  1.11 55
Average 1.07 53
Standard deviation 0.040 2.0
%RSD 3.8 3.8
Now admittedly this is an extreme case. With one operator getting recoveries of 97% and a second operator getting recoveries of 53%, I would suspect something is wrong. However, that doesn’t change the statistical evaluation. If I use %RSD for each operator, then it appears that the variability of the data with operator B (%RSD = 3.8) is much greater than the variability of operator A (%RSD = 2.0). However, when one looks at the data itself, and specifically at the standard deviations, the standard deviation in each case is the same, suggesting the variability in each case is the same.
If it doesn’t make sense to use %RSD as a measure of variability, what can we use instead? If we look at the data in Table A and Table B, perhaps we fall back to using just the standard deviation of the values themselves as a measure of variability. However, while that works in those two specific situations, what will happen in a significantly different case? Let’s suppose the 100% recovery value is not 2.03 µg/cm2, but is 20.3 µg/cm2. Table C has one possible data set for that situation.
Table C Data
  Data values % Recovery Values
Replicate 1 value 19.3 95
Replicate 2 value 19.6 97
Replicate 3 value 20.1 99
Average 19.7 97
Standard deviation 0.40 2.0
%RSD 2.1 2.1
Note that the 100% value in Table C is 10 times higher than the 100% value in Table A. In addition, the data for the three replicates are 10 times higher than the data in Table A. In this case, we look at the actual standard deviation values, we see that they are significantly different (0.040 for Table A vs. 0.40 for Table C), and we conclude that perhaps in this situation the %RSD is a better indication of variability of the data.
What we are faced with is a dilemma. What gives a better indication of variability of replicates, the standard deviation values (which appear a better measure in comparing A and B), or the %RSD values (which appear to give a better measure in comparing A and C).
One possible solution is to base the measure of variability on the standard deviation of the data itself as a percentage of the 100% recovery value. What this does is normalize the data, so that the variation in Table A and Table B are the same, but also the variation in the data in Table A and Table C are the same. The data expressing the standard deviation as a percentage of the 100% value is given below for each of the three cases covered above:
Table A Case: 100(0.040/2.03) = 2.0%
Table B Case: 100(0.040/2.03) = 2.0%
Table C Case: 100(0.40/20.3) = 2.0%
In looking at the data in the three situations, the variability of each operator seems to be the same. This measure (the standard deviation of the values as percentage of the 100% recovery value) appears to reflect that similarity. Furthermore, in cases where the data values might be more divergent, it would also appropriately reflect that greater variability.
Where does that leave us? Am I expecting people to start using this measure of variability in place of the %RSD for sampling recovery studies? Probably not. Part of the reason is that a variability criterion for sampling recovery studies is typically set at a relatively high level (15%-20% RSD), reflecting the high variability of recovery studies. Furthermore, if percent recoveries are relatively high (>80%), the difference between the proposed new measure and %RSD is somewhat minor. Finally, if percent recoveries are low but still acceptable (such as 50-65%), then the %RSD measurement will give a higher measure of variability, thus reflecting a worst case. So, while this proposed measure may provide a more scientific basis for the degree of variability, the existing method is not terribly wrong (particularly for something as variable as a swab recovery study).

Statistics for Visual Limits

This Cleaning Memo is an evaluation of a recent Pharmaceutical Technology article entitled “Statistically Justifiable Visible Residue Limits”, by M. Ovais (March 2010 issue, pages 58-71). The author asserts that “current methods for establishing visible residue limits are not statistically justifiable”. The author presents an example of determining a “visual residue limit” by spiking studies. In a spiking study, coupons are spiked at different levels, and a panel of observers looks at each panel under defined viewing conditions to determine the nature of the residue. Without going into the detail, the author provides a “logistic-regression” model to determine the “probability of detection” of the residue at that selected level. Needless to say, what this results in is a higher limit (a worst-case) than what would be determined by a consensus of multiple observers.
There are several questions to ask. Is this statistically correct? And, is this statistical evaluation really necessary? I can’t answer the first question; I’ll leave that up to the statisticians. While one answer to the second question is “you can certainly do it because it results in a higher visual limit” (a higher visual limit, contrary to what is often thought, is actually a worst-case), I’ll give my answer to the second question below.
However, to do that it is necessary to clarify a few things. One is that many publications (apparently including this Pharm Tech article) list the “visual limit” as the lowest spiked level at which observers can consistently see any residue on the spiked surface (of course this is under defined viewing conditions and for a defined residue and a defined surface, but that will be assumed throughout this discussion). In other words, if I spike a surface of 25 cm2 at different levels, then the spiked level at which all observers see even a speck on the spiked surface is the visual limit. While that may be one definition of visual limit, it is not a useful definition for cleaning validation purposes.
Why do I say it is not useful? The main reason is that the purpose of a visual limit is to say any surface viewed (under the same viewing conditions) that is visually clean has residue below that defined visual limit. Unfortunately, doing spiking studies and determining the lowest spiked level which has any residue on the spiked surface can’t be used in that way. Why? First, remember that the worst case for a visual limit is a high value, not a low value. Defining the visual limit in this way presents an artificially low visual limit, which will allow one to state that the residue is below the specified value without a sound scientific basis.
The issue here is that when I spike at a fixed level (let’s say 1.0 μg/cm2), and only see a small speck of residue in a corner of the area spiked, I cannot really say that any surface that is visually clean has a residue level of less than 1.0 μg/cm2. If I spiked at 1.0 μg/cm2, and the surface was evenly covered (an ideal situation), then it would appropriate to say that the spiked surface truly represents 1.0 μg/cm2, and therefore any surface which is visually clean has a residue level below 1.0 μg/cm2.
What happens in the real world when I do spiking studies is that the residue is not evenly spread over the spiked surface. Instead, due to the drying effects (difference in drying between the edges of the spiked residue solution and the center of the spiked residue solution), I will typically see a “donut hole” effect, with differing amounts of residues on different parts of the spiked area. Therefore, if I spike at 1.0 μg/cm2, it is possible that some portions of the spiked area may have concentrations of 0.8 μg/cm2, while other portions have residue levels of 1.2 μg/cm2. Perhaps I can see the residue in the spiked areas where the surface concentration is 1.2 μg/cm2, but not see it at a surface concentration of 0.8 or 1.0 μg/cm2. In that case, I will say the visual limit is 1.0 μg/cm2, which would be misleading.
The “correct” way (or at least one correct way) to determine the visual limit is slightly different. The same spiking coupons are prepared. However, the visual limit is then defined as the lowest concentration (in μg/cm2) in which the entire spiked area is visually dirty or soiled (that is, the lowest level at which residue is seen across the entire spiked area). Defined in this way, there is then a scientific or rational justification for saying any surface observed that is visually clean has a residue below the spiked level. Note that in this case, there may be (or better stated, there will be) variations in amounts of residue on different parts of the spiked coupon. That is inevitable, because of the drying effect. However, this approach is one that should be used (and not the approach of defining the visual limit based on the lowest spiked level where any residue is seen).
Why am I explaining this to address a statistical question? First, there is a certain level of “safety” already built into the determination of the visual limit. When I spike at 1.0 μg/cm2, and state that the visual limit is 1.0 μg/cm2, the true visual limit is lower than that (due to the drying effects mentioned above). How much lower, I can’t say for certain; however, I suspect that the “true” visual limit is probably 0.8 μg/cm2 or below.
Secondly, defining a visual limit is not necessarily an exercise where I need to get that visual limit as low as possible. My preference in using “visually clean alone” is not to do a series of coupons spiked at different levels. I prefer to first calculate the residue limit (using traditional maximum allowable carryover calculations, for example) to determine the limit per surface area (for those of you who follow my writings, this is the L3 limit in μg/cm2). If my calculated residue limit is 4.0 μg/cm2, why do I need to establish a visual limit that may be as low as 1.0 μg/cm2? In this situation, I prefer first do a spiking study at 4.0 μg/cm2. If at that spiked level all observers were not able to see residue across the entire spiked area, then who cares what the visual limit is? I clearly cannot use visually clean alone in a protocol to establish that the residue is below the calculated limit. On the other hand, if I spike at 4.0 μg/cm2 and can readily see residue across the entire spiked area (albeit uneven amounts on different portions of the spiked area), then I have a rationale for saying that surfaces observed that are visually clean are, in fact, below the calculated limit.
Note that this last situation (of spiking at 4.0 μg/cm2) already has some (undefined) safety margin in that the “true” visual limit is somewhat lower (because of the uneven concentrations across the spiked surface). That said, my preference is to add an extra margin of safety. If the calculated residue limit is 4.0 μg/cm2, my preference is to spike at an additional lower level, such as 3.0 or 3.5 μg/cm2. If I can see residue across the entire spiked area at those lower levels, I have an additional margin of safety in my determination of a visual limit.
If what I have discussed is the proper way to implement determinations of visual limits for a use of ‘visually clean alone” (that is, without swab or rinse sampling), then it would appear that there are significant safety margins built into the evaluation, and that a statistical evaluation of the “probability of detection” may be nice to have, but is not necessary.
This leads me to my last point. In that same issue of Pharm Tech was a short article by Lynn Torbeck (my favorite statistician, as revealed in last month’s Cleaning Memo) entitled “The Role of Statistical Tests”. In it, Mr. Torbeck points out that statistical significance tests should only be used after one first determines that there is a practical difference between two data sets. If there is no practical difference, don’t perform the statistical tests.
While the published article on statistics for visual limits is not strictly on statistical significance, it does invoke statistical principles to determine whether future observers would get the same result (or better said, to set a limit such that there is a higher probability that future observers would also report the same visual limit). I would put forth that with the determination of visual limits properly done (that is, defining the visual limit as the lowest spiked level where all observers see residue across the entire spiked area) has sufficient safety margins (either inherent in the process or which can be added to the process) such that extensive statistical analysis adds little or no value.
Note that it is certainly possible to use the statistical approach to further make the visual limit higher (which is a worst case). However, I think a good understanding of what is involved in determining visual limits suggests that there are practical safeguards already built into the visual limit determination.

Visually Clean and Visual Limits

First, let’s clarify the purpose of a visual limit (VL) determination. VL is typically expressed in units of mass per surface area, such as µg/cm2.The purpose is to define a level at which a defined residue is clearly visible on a defined surface (typically a certain material of construction and surface roughness) under defined viewing conditions (typically distance, lighting and angle). The VL is then used to in this way: If the same surface is viewed under identical or more stringent viewing conditions and is visually clean, then the residue is present at a level below the VL. If the VL is equal to or below the limit established by a carryover calculation in a cleaning validation protocol, then the viewed surface meets the defined acceptance criterion without the need to perform swab sampling.
The relevant section in PIC/S PI 006-03 is section 7.11.3, which states:
Carry-over of product residues should meet defined criteria, for example the most stringent of the following three criteria:
(a) No more than 0.1% of the normal therapeutic dose of any product will appear in the maximum daily dose of the following product,
(b) No more than 10 ppm of any product will appear in another product,
(c) No quantity of residue should be visible on the equipment after cleaning procedures are performed. Spiking studies should determine the concentration at which most active ingredients are visible[emphasis added]
The question that we will address is whether for any use of a visually clean criterion, must I do spiking studies to determine the VL? In other words, most people agree that if I am using visually clean alone without any swab or rinse sampling for that surface, then I should do spiking studies to determine what the VL is. While it is commonly stated that visual limits are on the order of 1-4 µg/cm2, we all realize this variable. And further, if the calculated carryover limit is 0.1 µg/cm2, it is not likely that I will be able to use visually clean alone. Furthermore, if the calculated limit is above a certain value (which will depend on the viewing conditions), I would make the case that spiking studies are not required. For example, if the carryover limit were 13.5 µg/cm2 and a stainless steel surface could be viewed in a short distance (such as 2 feet) under reasonable lighting conditions, I would make case that the VL would be significantly below the calculated limit and spiking studies were not needed. Obviously, some kind of reasonableness needs to be used for this latter case; it might not apply if it were a white residue on a PTFE surface.
But the key question for this Cleaning Memo is whether I should (or am required to) determine the visual limit (by performing spiking studies) if I am only using visual examination to supplement swab and/or rinse sampling for the same examined surface.
There are at least two possible interpretations of the “recommendation” in PI 006-03 in Section 7.11.3.c that “Spiking studies should determine the concentration at which most active ingredients are visible”. One interpretation is that this requirement must be read in context, and that context comes from the phrase “the most stringent of the following three criteria”. That is, the requirement for spiking studies only applies if visually clean is the most stringent of the three criteria. If this is the case (and this is my interpretation), then current industry practices are generally consistent with this interpretation.
A second interpretation is that I must perform spiking studies (to determine the VL) for all cases where I use visually clean as a criterion in a cleaning validation protocol. Does this interpretation make scientific sense? Is there a scientific rationale why this should be done in all cases, and in specific where visual examination is done to supplement swab and/or rinse sampling? Let’s see what value it adds in the latter situation. To consider that, we’ll take a look at several examples.
Let’s suppose for the first example that we are in a situation where the calculated limit for a residue is below the visual limit. For example, suppose the calculated carryover limit is 1 µg/cm2 and the VL (if I were to do spiking studies) is 3 µg/cm2. In my cleaning validation protocol, I measure residues for a given surface by swabbing and get results below 1 µg/cm2. In addition, equipment is visually clean. I pass both those acceptance criteria. But, would a spiking study to actually determine the value of VL make any difference in whether I pass or fail. Yes, it might be “nice to know” that the VL for the residue is actually 3 µg/cm2, but is it necessary? My answer is “No”.
For this same example (the calculated limit for a residue is below the visual limit), let’s suppose that when I measure residues by swabbing the results are below 1 µg/cm2. But, my visual examination shows that the equipment is not visually clean. With those results, I would fail my protocol. Obviously it should be clear that the visual failure is not caused by the target residue (for example, the active), because I have analytical data showing that the active is at an acceptable level. The visual failure is most likely caused by other residues (such excipients and/or cleaning agents). Again, knowing a specific VL for the active offers no additional information or benefit.
There are two other situations in this same example (the calculated limit for a residue is below the visual limit). One is where the swab analytical data is above the acceptance limit and the equipment is visually clean. Another is where the swab analytical data is above the acceptance limit and the equipment is not visually clean. I won’t go into detail, for these two cases, but it should be clear in these situation that spiking studies to determine a value for the VL offers no value.
These last three paragraphs have deal with the example where the calculated limit for a residue is below the visual limit. Now we’ll consider the reverse situation, where the calculated limit for a residue is above the visual limit. For example, suppose the calculated carryover limit is 5 µg/cm2 and the VL (if I were to do spiking studies) is 3 µg/cm2. In my cleaning validation protocol, I measure residues for a given surface by swabbing and get results below 5 µg/cm2, and the equipment is visually clean. Does the fact that I have an experimentally determined VL add anything? Well, you might say that if I did a spiking study, I would know that the amount of the active was below 3 µg/cm2. But what is the value of knowing that? If for some reason, I measured the residue by swabbing and the result was 4 µg/cm2 (thus meeting the analytical limit), then if the VL was 3 µg/cm2, I would fail the visually clean criterion even though I did not perform a spiking study to determine the VL. In this situation, it is not possible to have a situation in which I failed the analytical swabbing limit but passed the visual limit.
To sum it up, in these examples (where I am both measuring the residue by swabbing to compare it to the calculated carryover limit and determining the equipment is visually clean), it is not the case that determining the VL by spiking studies adds any significant benefit to confirming that the sampled surfaces are acceptably clean.

More Uses for Visual Limit Determination

Last month I discussed that in a cleaning validation protocol, it is only required that one determines a Visual Limit (VL) by performing a spiking study if one is exclusively using visual examination as the acceptance criterion for defined equipment surfaces. However, there are other situations broadly under the category of cleaning validation, but not part of a cleaning validation (or qualification) protocol, where determining a Visual Limit may add value.
The first situation is where I am in the early stages of cleaning process development, and I want to know how effective my cleaning process is. In that situation I may not have an analytical method (and associated sampling recovery studies) developed and validated. However, if I can determine the carryover limit (in µg/cm2), I could readily either determine the lowest practical Visual Limit or else determine whether the residue spiked at the calculated limit was readily visible on spiked surfaces. In this way, I could determine whether the cleaning process was effective in terms of meeting the required residue acceptance limit. Note in this case, I would prefer to either determine the lowest VL or a VL at least 50% of the calculated limit in order to be convinced that the cleaning process was robust.
A second situation is using visual examination as a primary routine monitoring tool for the cleaning process after it has been validated. In this case, I would want to establish the VL at either the acceptance limit or the lowest possible Visual Limit. The purpose here is just to use this as a confirmation that the cleaning process is continuing to be effective after completion of the validation protocols. Of course, the assumption here is that all critical surfaces during the monitoring process can be inspected visually. Particularly for equipment cleaned by a CIP process, the level of visual inspection for routine monitoring may not be to the same degree as visual inspection during the validation protocols (where there may be significant disassembly and/or tank entry for visual inspection). In addition, if this monitoring is to be comprehensive, areas like pipes (that may not be inspected visually) should be sampled by rinse water testing to confirm acceptable monitoring results. Where this second situation may have particular value is in manual cleaning, where in many cases all critical surfaces are readily accessible for visual inspection.
It is important to understand what is being said here. I am not saying that a VL needs to be established for routine monitoring purposes. What I am saying is that if a VL is determined experimentally, then routine visual monitoring of equipment on every cleaning event contains a higher level of assurance that the cleaning process is acceptable. However, it may not have the same high level of confidence that may be present in the validation runs unless all critical surfaces are visually examined during the routine monitoring process.
A third situation for the use of visual limits is in an investigation where I have identified a possible cleaning process problem, but done so only after the cleaned equipment has been used for manufacture of another product. The key assumption for this use is that the equipment was visually examined after the problematic cleaning process. One way to check for the effect of the problematic cleaning process on the subsequently manufactured product is to take samples of that subsequent product and analyze it for residues that might have been left on equipment surfaces. This is possible, but not necessarily an easy task, because analyzing for residues of the prior active (and of the cleaning agent) in the subsequently manufactured product may require significant analytical method development. Assuming the validated methods are HPLC methods, these methods will have to take into consideration possible new interfering substances from that subsequently manufactured product. If the validated method is TOC, then it will be impossible to measure residues in the next product (assuming the next product is not just inorganics).
In this situation, if (as assumed) I have a visually examined the equipment after the problematic run and if I have a VL for a given residue, I can determine whether that residue was at an acceptable level. Note that I might want to do this for both the active ingredient (API) and the cleaning agent. In this case, however, I am allowed to recalculate the residue limit with the actual cleaned product (“Product A”) and the actual subsequent product as the next product (“Product B”) in the carryover calculation. I do not necessarily have to use the carryover calculation based on any worst case assumptions; using the actual two products is acceptable, and may result in a higher limit (thus making it more likely that the residues would be acceptable). In this situation, I would still treat this as a process deviation; however, the visual examination may help provide assurance that the subsequently manufactured product was acceptable. This would be part of my corrective action; I would still have to deal with preventive actions to keep whatever might have gone wrong from happening again.
These are just three possible used of visual examination apart from use in cleaning validation protocols. They certainly are not mandatory uses, but rather can be considered as part of risk assessment in designing and implementing an overall program.

Swab Sampling Recovery as a Function of Residue Level

I have generally taught that the percent swab sampling recovery decreases with increasing spiked level of residue, other things being equal. In other words, recovery at a level of X µg/cm2 should be higher than recovery at a level of 2X µg/cm2. I have not said how much higher the recovery percentage at the lower spiked level might be, but I believe that the difference based on spiked levels that differ by a factor of 2 or 3 would be minor, certainly compared to the variation of recovery percentages that are achieved by different operators or by the same operator on different days.
My rationale or explanation for such a belief is to present the analogy of using a snow shovel to pick up snow on a sidewalk. If I have one pass across the sidewalk to pick up as much snow as possible, it is likely that with a level of snow of only seven centimeters on the sidewalk, my one pass (one shovelful) might pick up a relatively large amount of snow in the “sampled” area. That value might be 60% to 80% of the snow present on the “sampled” area. On the other hand, if I were to use the same shovel and procedure on a sidewalk containing 70 centimeters of snow, in one shovelful I might get only 30% to 40% of the snow on the sidewalk. If snow and snow shovels are a foreign concept to you, you might translate the analogy into sand on a sidewalk and using a sand shovel.
It may be possible to take this analogy too far. If I were to propose only a layer of snow as thick as 0.1 millimeter, then the percentage of snow picked up by the shovel might be very low because the shovel would pass on top of the snow layer. However, I don’t believe this situation applies to sampling in cleaning validation protocols.
A second way to explain a decreasing recovery percentage with increasing spiked residue is to appeal to a solubility analogy. If the only mechanism of removal of the residue in the swabbing procedure (don’t get me wrong here; it probably is not the only removal mechanism), then it should take longer to dissolve a large amount of residue on the surface. That should result in recovery percentages decreasing as the spiked residue amount increased (again, other things being equal).
A third way to explain the situation is to appeal to saturation of the swab. Again, other things being equal, as the amount of residue spiked increases, it is more likely that the swab will be saturated with residue.
Now all that said, a recent publication (referred to in last month’s Cleaning Memo), “A Risk Management Approach to Cleaning Validation” by B W Pack and J D Hofer (Pharmaceutical Technology, Vol. 34, No. 6, June 2010), came to the opposite conclusion based on their data. The authors specifically state that “the predominant trend was that the average recovery of a compound increased as the spiked amount increased on a given material of construction.” Let me make it clear that the primary objective of this study was not to evaluate the effect of recovery as a function of spiked amount. This conclusion just appeared as an “offhand” comment in a discussion of “materials of construction”.
Here is a close approximation of a summary of their data this conclusion is based upon, which is for 316L stainless steel. They presented results for two compounds, Compound A (less soluble and more difficult to clean) and Compound B (more soluble and easier to clean). Note, however, when they made their conclusion about the “predominant trend”, only the data for Compound B was discussed immediately to support that trend.
Spike level % Recovery A % Recovery B
0.5 µg/swab ~53 74
5 µg/swab ~80 90
50 µg/swab ~95 95
On initial look, this data seems to support the conclusion that recovery increases with increasing spiked level. On the other hand, could there be other factors accounting for these differences, such as performance on different days (the same analyst did the swabbing, so that probably is not a factor unless swabbing skill changed over time). In addition, there was significantly more variation at the lowest spiked level. For example, the data for Compound A at the 0.5 µg/swab level varied from a low of about 15% to a high of about 77%. The data for Compound A at the 5 µg/swab level varied from a low of about 73% to a high of about 87%. The data for Compound A at the 50 µg/swab level varied from a low of about 91% to a high of about 99%. Is this a reflection of variability in the analytical method at low levels, or is it a function of variability of removal of residue from the surface at lower levels?
One other data set reported in the publication caused me concern. A strategy was presented for introducing new materials of construction into the grouping program. That involved comparison of the data for the new material of construction with data run at the same time for a control. In the example given, the one control was 316L stainless with compound A at 5.0 µg/swab (note that the publication lists the spiked amount as 5 µg/in2; however, that must be a typo since that level was not part of the original study). In any case, the recovery at this level was consistently at 98% (with a range of 95% to 101%). My point is, that if this residue level is the same as reported in the original study, there is a significant difference between the 98% reported in the additional study as compared to the approximately 80% reported in the original study.
In any case, the Pack and Hofer publication caused me to go back through my files for previous publications where there might be recovery data as a function of spiked level.
One publication was by P. Yang et al (Method Development of Swab Sampling for Cleaning Validation of a Residual Active Pharmaceutical Ingredient, Pharmaceutical Technology, January 2005, pp. 84-92). In this publication, recovery was done for an active on a nylon surface at two levels, 2.8 µg/swab and 4.0 µg/swab. The recovery was 80.4% at the lower level and 84.7 at the higher level. While this is consistent with the idea of increasing recovery with increasing spiked level, the difference between the two results is not practically significant to come to a conclusion.
Another publication is by S. Lombardo et al (Development of Surface Swabbing Procedures for a Cleaning Validation Program in a Biopharmaceutical Manufacturing Facility, Biotechnology and Bioengineering, December 5, 1995, pp. 513-519). In Figures 4(a) and 5(a) in this paper, a linear relation was shown between the amount of residue spiked onto the coupon and the amount recovered. It should be noted that in most situations, the data was based on two spiked levels, but the curve was forced through “zero”, which in essence gave three data points. The authors state that “a linear relationship prevails between the observed and anticipated contaminant values over the range investigated.” The fact that the curves were linear suggests that the recoveries were essentially the same at different spiked levels (the slope of the straight line should give the recovery expressed as a decimal). In this study the range of spiked levels differed by no more than a factor of about three
Another study was by K. Bader et al (Online Total Organic (TOC) as a Process Analytical Technology for Cleaning Validation Risk Management, Pharmaceutical Engineering, January/February 2009). In this publication, Figure 3 and Figure 4 are curves of recovered residue as compared to the positive control (the amount representing 100% recovery). Note that the two figures are the same data; however, Figure 3 presents individual swab technician results and Figure 4 presents the aggregate data. Although no conclusion is drawn from the data about the effect of spiked amount on percent recovery, the linear relationships suggests that percent recovery is the same over the evaluated range. For this study, the range from the low spiked level to the high spiked level was a factor of about five (5).
Another publication is by C. Glover (Validation of the Total Organic Carbon (TOC) Swab Sampling and Test Method, Journal of Pharmaceutical Science and Technology, September-October 2006, pp. 284-290). In Table I, percent recovery is given as function of five spiked levels from 5 µg to 100 µg (a range representing a factor of 20). The reported data showed a general decrease in percent recovery as a function of increasing spiked amount. However, the data at the lower levels gave recoveries of greater than 150%, while the recoveries at the higher levels were close to 100%. That data suggests some issues with the TOC analysis.
A final publication is M. A. Strege et al (Total Organic Carbon Analysis of Swab Samples for the Cleaning Validation of Bioprocess Fermentation Equipment, BioPharm International, April 1996). In this publication, Table 2 lists the percent recoveries from stainless steel for three dilutions of a fermentation cell paste. In this situation, the recovery increased with increasing spike level from a low of 75% to a high of 103%. This involved a range with a factor of four between the top and bottom spiked levels.
So, where does this leave us? The data from published studies sometimes show increasing percent recoveries with increasing spiked amounts, sometimes show no change in percent recoveries with increasing spiked amount, and sometimes show decreasing recoveries with increasing spiked amount. I should point out that in none of these studies cited was the stated objective to determine the relationship between the percent recovery and the amount of residue spiked.
If anyone has published any studies that can help elucidate this issue, I’d like to hear from them. If anyone would like to perform a study to specifically evaluate the relationship between percent recovery and spiked amount, I would be more than happy to assist in the design of it so that appropriate conclusions can be drawn. My only caution is that it would be best to avoid using TOC as the analytical method because of control of the sources of TOC. Furthermore, randomization of sampling order must be considered.
Until such time as a definitive study is published and confirmed, it would probably be best to stick with my original contention (based on a common sense understanding of what happens in a swabbing procedure) that other things being equal, percent recovery decreases with increasing spiked level, without stating how significant that decrease might be.

Understanding the Cleaning Process in 2010

In January 2005 I wrote a Cleaning Memo entitled “Understanding the Cleaning Process”. That Cleaning Memo was in response to the FDA report on risk-based approaches to pharmaceutical CGMPs. At the end of that Cleaning Memo, I encouraged manufacturers to “explore more fully what is occurring in a cleaning process, and then to use that knowledge to design more effective and more efficient cleaning processes, as well as simpler ways to validate those processes.” That encouragement is even more important now, based on the 2008 FDA draft process validation guidance (which reportedly will be finalized in the first quarter of 2011) that defines “design and development” as the first stage of the validation process.
Understanding what is happening in the cleaning process (which also includes what is happening in the equipment soiling process) is a key to using these new principles to more effectively and efficiently implement cleaning validation. One of the least quoted sections of the 1993 FDA cleaning validation guidance is in Section IV (“Evaluation of Cleaning Validation”) where the statement is made that “Answers to these questions may also identify steps that can be eliminated for more effective measures and result in resource savings for the company.” This statement, in which the FDA is suggesting “resource saving”, is made in the context of question such as “… at what point does a piece of equipment or system become clean? Does it have to be scrubbed by hand? What is accomplished by hand scrubbing rather than just a solvent wash? How variable are manual cleaning processes from batch to batch and product to product?” Now don’t think for a minute that the FDA wants “resource savings” so that your firm can be more profitable. I don’t know what the FDA’s reason for this statement was in 1993, but in 2010 the reason is clearly that the FDA is concerned about costs of drugs to consumers.
So, when the FDA is encouraging the industry to look for “resource savings”, why are some of us still doing things the way we did in 1993 when we had more limited information of what was happening in our cleaning processes? What are some things that we can do based on a better understanding of the cleaning process?
Well, one example is the ISPE RiskMaPP approach to dealing with highly hazardous drug actives (such as actives that might be genotoxic, mutagenic, or teratogenic). While I criticized (in my November 2010 Cleaning Memo) the RiskMaPP document for the way it critiqued previous methods of setting limits, the fundamental approach in that RiskMaPP document of setting health-based limits for highly hazardous actives (as opposed to previous approaches of using dedicated equipment or requiring that limits be set as non-detectable by the best available analytical technique) is an good example of using knowledge of the cleaning process to enable these highly hazardous actives to be manufactured in the same facility or on the same equipment (with appropriate controls in place).
Another example is understanding what is happening during the dirty hold time (DHT). The traditional approach has been to require a challenge of the maximum time during a validation protocol. However, it is clear (at least in some cases) that an increase in time does not change the difficulty of cleaning. In other cases, the difficulty of cleaning may increase with time up to a certain point (for example, when a liquid is “dry”), and then not change after that. An evaluation of how to approach the DHT will depend on our understanding of how the diffculty of cleaning might change over time. That means not only understanding the nature of drying, but also whether bioburden proliferates during the DHT and whether degradation of the active is accelerated during the DHT. Of course, the approach now should be to understand those factors, and design the cleaning process with those factors in mind, such that the challenges are addressed during the design/development stage rather than during the qualification protocol.
A third example is dealing with campaign length, where so-called minor cleaning (such as vacuuming or a water flush) is performed between batches, and a validated cleaning process is only performed at the end of a campaign. The question comes up, what if the campaign is sometimes 5 batches and sometime 7 batches? Do my qualification protocols have to be at the maximum campaign length? Well, absent any information on the effect of campaign length on the diffculty of cleaning, it makes sense to perform a qualification protocol at the end of the maximum of 7 batches. If only production scheduling would cooperate by providing those number of batches for the required number of validation runs, it might be easy. However, scheduling is not always that nice; plus there might be a time when they want to run 8 batches in a campaign. What can be done?
Well, the secret phrase in the above paragraph was “absent any information on the effect of campaign length on the diffculty of cleaning”. Is there information or data I can obtain from laboratory or developmental studies, or from “sufficiently similar” products or processes, that might allow me to determine the effect of campaign length on diffculty of cleaning? Particularly if I can demonstrate that campaign length has no effect on difficulty of cleaning, performing my qualification protocols after just one batch may be adequate. Again, this may also involve determining information such as bioburden proliferation and/or degradation of the active as a function of campaign length.
There are other approaches to understanding the cleaning process which were discussed in my April 2008 Cleaning Memo (“What Have We Learned in the Last Two Decades?”) that also can be considered. The issue is this – if a pharmaceutical manufacturer is to thrive in the coming decades, the approach of new drugs with significant advantages is still the primary objective. However, low cost (but meeting current CGMP) production should be a secondary, but vital objective.

A Critique of Cleaning Validation Issues in ISPE’s RiskMaPP

ISPE has issued the document “Risk-Based Manufacture of Pharmaceutical Products” (Volume 7 of their Baseline Guides, September 2010). As I understand it, the effort to write this guide started with concerns over regulatory bodies tending to require dedicated equipment and/or facilities for certain highly hazardous drug actives, such as potent drugs, hormones, genotoxic compounds, and cytotoxic compounds. The major rationale for this ISPE guide was to counteract this approach by providing for an analysis of safety/toxicity data of these “highly hazardous” drug actives to determine a level that might be a negligible (but acceptable) risk in other drug products (thus allowing, with appropriate controls, the ability to manufacture in non-dedicated equipment/facilities).
This effort involved setting limits for these “highly hazardous” actives based on what is called a “health based limit”, that is a limit based on a toxicological evaluation of the relevant safety data (typically a No Observable Adverse Effect Level or NOAEL). This health-based limit is called ADE, or Acceptable Daily Exposure, in the ISPE guide. This effort for setting limits for cleaning purposes for these highly hazardous actives is to be applauded.
Unfortunately, the guide goes beyond that basic focus to discuss cleaning validation in general, and to imply that this method for cleaning validation is appropriate in all cases, including what I will call “conventional” actives that don’t have these highly hazardous concerns. Specifically, the guide states that current methods of setting limits such as “1/1,000th of the lowest clinical dose or 10 ppm in a batch” are “arbitrary limits”. (page 42) Furthermore, “the use of arbitrary non-health-based limits is not scientifically justified” if data exists to calculate a health-based limit. (page 45) Additionally, “Another non-science-based approach for setting cleaning limits is the use of the ‘10 ppm’ specification” (page 46), and “if default values such as 10 [ppm] are used arbitrarily to set allowable residue limits, they may be lower than they need to be from a health perspective” (page 45). It further states that using such values ignores toxicological data and can be “too restrictive or not sufficiently restrictive”. (page 42)
In other words, according to this guide, current methods of setting limits for cleaning validation purposes (which have been used for at least the last 17 years) are “arbitrary” and “not science-based”. I do not find those assertions to be factually correct, nor is there anything in the guide to support those assertions.
Part of the issue is that the ISPE guide sets up a “straw man” in terms of how limits are currently set. Criteria such as 1/1,000th dose or 10 ppm are critiqued individually. As used by manufacturers, the current method (note that this is what I consider the best approach; I am not suggesting that all pharmaceutical companies use this approach) involves setting limits based on the most stringent of these three criteria:
  • 1/1,000th of a dose of an active in a dose of the next drug product
  • 10 ppm of the active in a dose of the next drug product
  • Equipment surfaces are visually clean
In other words, these are not considered independently. Now, this works well for most drug actives. Does it work well for a teratogenic drug active? Of course not, and nobody with any significant experience in cleaning validation would state that limits are set on 1/1,000th of a dose for a teratogenic drug active. The relevant safety concern for the teratogenic drug active is not the pharmacologic effect; therefore the cleaning validation residue level is not set based on a fraction of the drug dosage. The general approach based on the PIC/S recommendations for cleaning validation is that the equipment either be dedicated or that a residue level be established as “non-detectable” by the best available analytical technique. [Note here that my recommendation for this “non-detectable” requirement has been to have a toxicologists confirm that the non-detectable level is an acceptable risk; this addresses “best available techniques” that aren’t good enough.]
It is latter concern (of dedicating equipment or establishing limits as non-detectable) that I would think that the ISPE guide would want to counteract by offering the idea of a toxicological evaluation to set an acceptable limit. I don’t see the value (or rationale) for the ISPE guide stating that current methods of setting limits are “arbitrary” and “nonscientific”.
Some of the statements made in the ISPE guide about the current way of setting are true. If limits are set at 1/1,000th of a dose, with some conventional drug actives the level of protection will be more than is necessarily required from a patient safety concern. This is to be expected with a “one size fits all” approach (but don’t get me wrong; this one size fits all doesn’t apply to active where the hazard is not related to the therapeutic effect). However, this does not mean that limits are set in a non-scientific way.
In the training session introducing the guide, one of the speakers stated that there is a degree of judgment in establishing appropriate factors for calculating health-based limits. Does that mean that the limits are non-scientific? Of course not! What it means is that different toxicologists may come up with different numbers for the health-based limits, one being more stringent than the other. An analogous situation exists with the 1/1,000th calculation; it is more stringent than it needs to be for some actives, but for all conventional actives it provides a safe limit.
If the health-based criterion is applied to conventional actives, the acceptable levels given in some example in the ISPE guide are extremely high, such that there might be other concerns other than patient safety. Those concerns might include stability, production efficiency, and interference with the bioavailability of the next active. For example, an ADE of an unspecified NSAID active is given as 40 mg/day. A daily dose of this NSAID is given as 800 mg. (page 101) The implication here is that 1/20th of the daily dose is a safe level to have in a daily dose of a subsequent product. If this ADE were present in a subsequently manufactured drug product involving 250 mg tablets given eight tablets per day, the acceptable concentration of that active in that next product would be 20,000 ppm (2%). Would anyone really consider allowing any drug active to be in a subsequent drug product at that level? It just wouldn’t be CGMP. In case you might think I am overstating the case, there is another example given for an “antisense” active, where the ADE is 0.5 mg/day and the daily dose is 10 mg. (page 102) The implication here is 1/20th of a dose of the antisense active is a safe level. Again, this might be safe from a toxicologist viewpoint, but I doubt if most companies would (or should) allow or permit such levels.
Furthermore, at those high levels, it is likely that residue on equipment surfaces (those that could be observed) would be visually dirty. Does it make sense to set limits to allow such a situation (unless for conventional drugs we really want to only require “visually clean” for surfaces which are readily observable)?
The reply to my objection might be given that I have not read the entire document, and that just because these are limits, it doesn’t mean that manufacturers want to allow residue nearing those levels. The concept of “Margin of Safety” is presented in the ISPE document (page 42). The “Margin of Safety” is not to be confused with so-called “safety factors” that are used in setting limits based on a fraction of the dose. The “Margin of Safety” as defined in the ISPE guide, is the difference or “distance” between the established limit and the actual residue data obtained in a cleaning validation protocol. The argument might be made that while the safe limit (as determined by the ADE) is relatively high, the actual data is much below that ADE-based limit., so the situations I discussed are not likely to happen.
My response to that reply is that those situations are allowed if the limit is that high. Yes, under either ADE limits or 1/1,000th limits, I would like my residue data in protocols to be significantly below my established limit. In other words, I want a robust cleaning process. But this is where the ISPE guide and I differ. The ISPE guide states that “Evaluation of the cleaning validation data is the only way to ensure that any residuals after cleaning are as low as possible below the health-based criteria….” [emphasis added] (page 43) It sounds like the guide, while arguing that limits in some cases are more stringent than they need to be, suggests that manufacturers should still clean to residue levels as low as possible. So that you don’t think I am taking this out of context, the guide also states that “Efforts should be made to ensure that cleaning procedures provide as large a safety margin as possible.” (page 45) In other words, the guide seems to be saying that for conventional actives, limits can be looser (that is, higher), but you should still clean to the same low level so that the “Margin of Safety” is as great as possible.
There are other concerns I have about the document. However, the main concern I have is the way the document inappropriately “trashes” current methods of setting limits as arbitrary and non-scientific. I think an easy fix could be made to the document by revising it to what I think is its original focus. That is, take out all general references to how limits are set, and focus on the main issue, which is that health-based limits are appropriate for dealing with drug actives that have considerable hazard concerns unrelated to the therapeutic effect. Those actives include hormones, potent steroids, and actives that are genotoxicity, cytotoxicity, or have reproductive hazards. The rationale for this suggested change is that the purpose of the ISPE guide seems to be having a science- based method for avoiding manufacture of such compounds in dedicated facilities/equipment. Supposedly there is in development another ISPE guide on cleaning validation. It would seem that that cleaning validation guide should be the one to specifically address the issue of whether limits for “conventional” actives should be changed.
Some of you might want to know why I am voicing my objections now. After all, the guide has been in development for five years. In the fall of 2007, I contacted ISPE about obtaining the draft document to provide comments, but found out that the comment period had just closed. However, I voiced my objections (essentially the same objections expressed here) back in 2008. These concerns were not specifically about the draft of the document at that time, but were based on presentations made by Andy Walsh (a RiskMaPP task force member) at an ISPE meeting in Washington DC in the summer of 2008. I gave a webinar entitled “Are we Setting Limits Correctly?” in August 2008 and wrote a Cleaning Memo with the same title in October 2008. However, whatever happened in the past, it should be clear that the published guide needs correction.
For clarification, I am not concerned about the guide’s approach to setting limits for highly hazardous actives. The approach of a toxicological evaluation based on those hazards is appropriate. If those evaluations result in making products in non-dedicated facilities/equipment, then that is a needed step forward. Note that this approach is consistent with the December 2009 EMA document EMA/INS/GMP/809387/2009,“Update on revision of Chapters 3 and 5 of the GMP Guide: Dedicated facilities"

Plasmids for Vaccine Validation

pCMV-S (also known as pRc/CMV-HBs) (Figure 5) is widely used to validate DNA vaccine delivery and formulation strategies. This plasmid expresses the hepatitis B surface antigen (HBsAg) under the control of the CMV immediate-early promoter.
pCMVHB-S2.S expresses the small and middle forms of recombinant HBsAg. The plasmid can be used to generate anti-HBsAg antibodies like pCMV-S. It is also used to fuse other sequences to the S form of HBsAg.
The antibody response to pCMV-S and/or pCMVHB-S2.S can be tested via ELISA using Aldevron's recombinant HBsAg. Both of these plasmids are available free of charge for research applications (Table 5). These vectors are covered and described by United States Patent 6,635,624 which is available at www.uspto.gov.
Figure 5 — Validate your DNA vaccine delivery and formulation with plasmids provided free, courtesy of Aldevron.

pRc CMV-HBsS