Last month I discussed that in a cleaning validation protocol, it is only required that one determines a Visual Limit (VL) by performing a spiking study if one is exclusively using visual examination as the acceptance criterion for defined equipment surfaces. However, there are other situations broadly under the category of cleaning validation, but not part of a cleaning validation (or qualification) protocol, where determining a Visual Limit may add value.
The first situation is where I am in the early stages of cleaning process development, and I want to know how effective my cleaning process is. In that situation I may not have an analytical method (and associated sampling recovery studies) developed and validated. However, if I can determine the carryover limit (in µg/cm2), I could readily either determine the lowest practical Visual Limit or else determine whether the residue spiked at the calculated limit was readily visible on spiked surfaces. In this way, I could determine whether the cleaning process was effective in terms of meeting the required residue acceptance limit. Note in this case, I would prefer to either determine the lowest VL or a VL at least 50% of the calculated limit in order to be convinced that the cleaning process was robust.
A second situation is using visual examination as a primary routine monitoring tool for the cleaning process after it has been validated. In this case, I would want to establish the VL at either the acceptance limit or the lowest possible Visual Limit. The purpose here is just to use this as a confirmation that the cleaning process is continuing to be effective after completion of the validation protocols. Of course, the assumption here is that all critical surfaces during the monitoring process can be inspected visually. Particularly for equipment cleaned by a CIP process, the level of visual inspection for routine monitoring may not be to the same degree as visual inspection during the validation protocols (where there may be significant disassembly and/or tank entry for visual inspection). In addition, if this monitoring is to be comprehensive, areas like pipes (that may not be inspected visually) should be sampled by rinse water testing to confirm acceptable monitoring results. Where this second situation may have particular value is in manual cleaning, where in many cases all critical surfaces are readily accessible for visual inspection.
It is important to understand what is being said here. I am not saying that a VL needs to be established for routine monitoring purposes. What I am saying is that if a VL is determined experimentally, then routine visual monitoring of equipment on every cleaning event contains a higher level of assurance that the cleaning process is acceptable. However, it may not have the same high level of confidence that may be present in the validation runs unless all critical surfaces are visually examined during the routine monitoring process.
A third situation for the use of visual limits is in an investigation where I have identified a possible cleaning process problem, but done so only after the cleaned equipment has been used for manufacture of another product. The key assumption for this use is that the equipment was visually examined after the problematic cleaning process. One way to check for the effect of the problematic cleaning process on the subsequently manufactured product is to take samples of that subsequent product and analyze it for residues that might have been left on equipment surfaces. This is possible, but not necessarily an easy task, because analyzing for residues of the prior active (and of the cleaning agent) in the subsequently manufactured product may require significant analytical method development. Assuming the validated methods are HPLC methods, these methods will have to take into consideration possible new interfering substances from that subsequently manufactured product. If the validated method is TOC, then it will be impossible to measure residues in the next product (assuming the next product is not just inorganics).
In this situation, if (as assumed) I have a visually examined the equipment after the problematic run and if I have a VL for a given residue, I can determine whether that residue was at an acceptable level. Note that I might want to do this for both the active ingredient (API) and the cleaning agent. In this case, however, I am allowed to recalculate the residue limit with the actual cleaned product (“Product A”) and the actual subsequent product as the next product (“Product B”) in the carryover calculation. I do not necessarily have to use the carryover calculation based on any worst case assumptions; using the actual two products is acceptable, and may result in a higher limit (thus making it more likely that the residues would be acceptable). In this situation, I would still treat this as a process deviation; however, the visual examination may help provide assurance that the subsequently manufactured product was acceptable. This would be part of my corrective action; I would still have to deal with preventive actions to keep whatever might have gone wrong from happening again.
These are just three possible used of visual examination apart from use in cleaning validation protocols. They certainly are not mandatory uses, but rather can be considered as part of risk assessment in designing and implementing an overall program.
Validation refers to establishing documented evidence that a process or system, when operated within established parameters, can perform effectively and reproducibly to produce a medicinal product meeting its predetermined specifications and quality attributes
Monday, December 6, 2010
Swab Sampling Recovery as a Function of Residue Level
I have generally taught that the percent swab sampling recovery decreases with increasing spiked level of residue, other things being equal. In other words, recovery at a level of X µg/cm2 should be higher than recovery at a level of 2X µg/cm2. I have not said how much higher the recovery percentage at the lower spiked level might be, but I believe that the difference based on spiked levels that differ by a factor of 2 or 3 would be minor, certainly compared to the variation of recovery percentages that are achieved by different operators or by the same operator on different days.
My rationale or explanation for such a belief is to present the analogy of using a snow shovel to pick up snow on a sidewalk. If I have one pass across the sidewalk to pick up as much snow as possible, it is likely that with a level of snow of only seven centimeters on the sidewalk, my one pass (one shovelful) might pick up a relatively large amount of snow in the “sampled” area. That value might be 60% to 80% of the snow present on the “sampled” area. On the other hand, if I were to use the same shovel and procedure on a sidewalk containing 70 centimeters of snow, in one shovelful I might get only 30% to 40% of the snow on the sidewalk. If snow and snow shovels are a foreign concept to you, you might translate the analogy into sand on a sidewalk and using a sand shovel.
It may be possible to take this analogy too far. If I were to propose only a layer of snow as thick as 0.1 millimeter, then the percentage of snow picked up by the shovel might be very low because the shovel would pass on top of the snow layer. However, I don’t believe this situation applies to sampling in cleaning validation protocols.
A second way to explain a decreasing recovery percentage with increasing spiked residue is to appeal to a solubility analogy. If the only mechanism of removal of the residue in the swabbing procedure (don’t get me wrong here; it probably is not the only removal mechanism), then it should take longer to dissolve a large amount of residue on the surface. That should result in recovery percentages decreasing as the spiked residue amount increased (again, other things being equal).
A third way to explain the situation is to appeal to saturation of the swab. Again, other things being equal, as the amount of residue spiked increases, it is more likely that the swab will be saturated with residue.
Now all that said, a recent publication (referred to in last month’s Cleaning Memo), “A Risk Management Approach to Cleaning Validation” by B W Pack and J D Hofer (Pharmaceutical Technology, Vol. 34, No. 6, June 2010), came to the opposite conclusion based on their data. The authors specifically state that “the predominant trend was that the average recovery of a compound increased as the spiked amount increased on a given material of construction.” Let me make it clear that the primary objective of this study was not to evaluate the effect of recovery as a function of spiked amount. This conclusion just appeared as an “offhand” comment in a discussion of “materials of construction”.
Here is a close approximation of a summary of their data this conclusion is based upon, which is for 316L stainless steel. They presented results for two compounds, Compound A (less soluble and more difficult to clean) and Compound B (more soluble and easier to clean). Note, however, when they made their conclusion about the “predominant trend”, only the data for Compound B was discussed immediately to support that trend.
On initial look, this data seems to support the conclusion that recovery increases with increasing spiked level. On the other hand, could there be other factors accounting for these differences, such as performance on different days (the same analyst did the swabbing, so that probably is not a factor unless swabbing skill changed over time). In addition, there was significantly more variation at the lowest spiked level. For example, the data for Compound A at the 0.5 µg/swab level varied from a low of about 15% to a high of about 77%. The data for Compound A at the 5 µg/swab level varied from a low of about 73% to a high of about 87%. The data for Compound A at the 50 µg/swab level varied from a low of about 91% to a high of about 99%. Is this a reflection of variability in the analytical method at low levels, or is it a function of variability of removal of residue from the surface at lower levels?
One other data set reported in the publication caused me concern. A strategy was presented for introducing new materials of construction into the grouping program. That involved comparison of the data for the new material of construction with data run at the same time for a control. In the example given, the one control was 316L stainless with compound A at 5.0 µg/swab (note that the publication lists the spiked amount as 5 µg/in2; however, that must be a typo since that level was not part of the original study). In any case, the recovery at this level was consistently at 98% (with a range of 95% to 101%). My point is, that if this residue level is the same as reported in the original study, there is a significant difference between the 98% reported in the additional study as compared to the approximately 80% reported in the original study.
In any case, the Pack and Hofer publication caused me to go back through my files for previous publications where there might be recovery data as a function of spiked level.
One publication was by P. Yang et al (Method Development of Swab Sampling for Cleaning Validation of a Residual Active Pharmaceutical Ingredient, Pharmaceutical Technology, January 2005, pp. 84-92). In this publication, recovery was done for an active on a nylon surface at two levels, 2.8 µg/swab and 4.0 µg/swab. The recovery was 80.4% at the lower level and 84.7 at the higher level. While this is consistent with the idea of increasing recovery with increasing spiked level, the difference between the two results is not practically significant to come to a conclusion.
Another publication is by S. Lombardo et al (Development of Surface Swabbing Procedures for a Cleaning Validation Program in a Biopharmaceutical Manufacturing Facility, Biotechnology and Bioengineering, December 5, 1995, pp. 513-519). In Figures 4(a) and 5(a) in this paper, a linear relation was shown between the amount of residue spiked onto the coupon and the amount recovered. It should be noted that in most situations, the data was based on two spiked levels, but the curve was forced through “zero”, which in essence gave three data points. The authors state that “a linear relationship prevails between the observed and anticipated contaminant values over the range investigated.” The fact that the curves were linear suggests that the recoveries were essentially the same at different spiked levels (the slope of the straight line should give the recovery expressed as a decimal). In this study the range of spiked levels differed by no more than a factor of about three
Another study was by K. Bader et al (Online Total Organic (TOC) as a Process Analytical Technology for Cleaning Validation Risk Management, Pharmaceutical Engineering, January/February 2009). In this publication, Figure 3 and Figure 4 are curves of recovered residue as compared to the positive control (the amount representing 100% recovery). Note that the two figures are the same data; however, Figure 3 presents individual swab technician results and Figure 4 presents the aggregate data. Although no conclusion is drawn from the data about the effect of spiked amount on percent recovery, the linear relationships suggests that percent recovery is the same over the evaluated range. For this study, the range from the low spiked level to the high spiked level was a factor of about five (5).
Another publication is by C. Glover (Validation of the Total Organic Carbon (TOC) Swab Sampling and Test Method, Journal of Pharmaceutical Science and Technology, September-October 2006, pp. 284-290). In Table I, percent recovery is given as function of five spiked levels from 5 µg to 100 µg (a range representing a factor of 20). The reported data showed a general decrease in percent recovery as a function of increasing spiked amount. However, the data at the lower levels gave recoveries of greater than 150%, while the recoveries at the higher levels were close to 100%. That data suggests some issues with the TOC analysis.
A final publication is M. A. Strege et al (Total Organic Carbon Analysis of Swab Samples for the Cleaning Validation of Bioprocess Fermentation Equipment, BioPharm International, April 1996). In this publication, Table 2 lists the percent recoveries from stainless steel for three dilutions of a fermentation cell paste. In this situation, the recovery increased with increasing spike level from a low of 75% to a high of 103%. This involved a range with a factor of four between the top and bottom spiked levels.
So, where does this leave us? The data from published studies sometimes show increasing percent recoveries with increasing spiked amounts, sometimes show no change in percent recoveries with increasing spiked amount, and sometimes show decreasing recoveries with increasing spiked amount. I should point out that in none of these studies cited was the stated objective to determine the relationship between the percent recovery and the amount of residue spiked.
If anyone has published any studies that can help elucidate this issue, I’d like to hear from them. If anyone would like to perform a study to specifically evaluate the relationship between percent recovery and spiked amount, I would be more than happy to assist in the design of it so that appropriate conclusions can be drawn. My only caution is that it would be best to avoid using TOC as the analytical method because of control of the sources of TOC. Furthermore, randomization of sampling order must be considered.
Until such time as a definitive study is published and confirmed, it would probably be best to stick with my original contention (based on a common sense understanding of what happens in a swabbing procedure) that other things being equal, percent recovery decreases with increasing spiked level, without stating how significant that decrease might be.
My rationale or explanation for such a belief is to present the analogy of using a snow shovel to pick up snow on a sidewalk. If I have one pass across the sidewalk to pick up as much snow as possible, it is likely that with a level of snow of only seven centimeters on the sidewalk, my one pass (one shovelful) might pick up a relatively large amount of snow in the “sampled” area. That value might be 60% to 80% of the snow present on the “sampled” area. On the other hand, if I were to use the same shovel and procedure on a sidewalk containing 70 centimeters of snow, in one shovelful I might get only 30% to 40% of the snow on the sidewalk. If snow and snow shovels are a foreign concept to you, you might translate the analogy into sand on a sidewalk and using a sand shovel.
It may be possible to take this analogy too far. If I were to propose only a layer of snow as thick as 0.1 millimeter, then the percentage of snow picked up by the shovel might be very low because the shovel would pass on top of the snow layer. However, I don’t believe this situation applies to sampling in cleaning validation protocols.
A second way to explain a decreasing recovery percentage with increasing spiked residue is to appeal to a solubility analogy. If the only mechanism of removal of the residue in the swabbing procedure (don’t get me wrong here; it probably is not the only removal mechanism), then it should take longer to dissolve a large amount of residue on the surface. That should result in recovery percentages decreasing as the spiked residue amount increased (again, other things being equal).
A third way to explain the situation is to appeal to saturation of the swab. Again, other things being equal, as the amount of residue spiked increases, it is more likely that the swab will be saturated with residue.
Now all that said, a recent publication (referred to in last month’s Cleaning Memo), “A Risk Management Approach to Cleaning Validation” by B W Pack and J D Hofer (Pharmaceutical Technology, Vol. 34, No. 6, June 2010), came to the opposite conclusion based on their data. The authors specifically state that “the predominant trend was that the average recovery of a compound increased as the spiked amount increased on a given material of construction.” Let me make it clear that the primary objective of this study was not to evaluate the effect of recovery as a function of spiked amount. This conclusion just appeared as an “offhand” comment in a discussion of “materials of construction”.
Here is a close approximation of a summary of their data this conclusion is based upon, which is for 316L stainless steel. They presented results for two compounds, Compound A (less soluble and more difficult to clean) and Compound B (more soluble and easier to clean). Note, however, when they made their conclusion about the “predominant trend”, only the data for Compound B was discussed immediately to support that trend.
| Spike level | % Recovery A | % Recovery B |
| 0.5 µg/swab | ~53 | 74 |
| 5 µg/swab | ~80 | 90 |
| 50 µg/swab | ~95 | 95 |
One other data set reported in the publication caused me concern. A strategy was presented for introducing new materials of construction into the grouping program. That involved comparison of the data for the new material of construction with data run at the same time for a control. In the example given, the one control was 316L stainless with compound A at 5.0 µg/swab (note that the publication lists the spiked amount as 5 µg/in2; however, that must be a typo since that level was not part of the original study). In any case, the recovery at this level was consistently at 98% (with a range of 95% to 101%). My point is, that if this residue level is the same as reported in the original study, there is a significant difference between the 98% reported in the additional study as compared to the approximately 80% reported in the original study.
In any case, the Pack and Hofer publication caused me to go back through my files for previous publications where there might be recovery data as a function of spiked level.
One publication was by P. Yang et al (Method Development of Swab Sampling for Cleaning Validation of a Residual Active Pharmaceutical Ingredient, Pharmaceutical Technology, January 2005, pp. 84-92). In this publication, recovery was done for an active on a nylon surface at two levels, 2.8 µg/swab and 4.0 µg/swab. The recovery was 80.4% at the lower level and 84.7 at the higher level. While this is consistent with the idea of increasing recovery with increasing spiked level, the difference between the two results is not practically significant to come to a conclusion.
Another publication is by S. Lombardo et al (Development of Surface Swabbing Procedures for a Cleaning Validation Program in a Biopharmaceutical Manufacturing Facility, Biotechnology and Bioengineering, December 5, 1995, pp. 513-519). In Figures 4(a) and 5(a) in this paper, a linear relation was shown between the amount of residue spiked onto the coupon and the amount recovered. It should be noted that in most situations, the data was based on two spiked levels, but the curve was forced through “zero”, which in essence gave three data points. The authors state that “a linear relationship prevails between the observed and anticipated contaminant values over the range investigated.” The fact that the curves were linear suggests that the recoveries were essentially the same at different spiked levels (the slope of the straight line should give the recovery expressed as a decimal). In this study the range of spiked levels differed by no more than a factor of about three
Another study was by K. Bader et al (Online Total Organic (TOC) as a Process Analytical Technology for Cleaning Validation Risk Management, Pharmaceutical Engineering, January/February 2009). In this publication, Figure 3 and Figure 4 are curves of recovered residue as compared to the positive control (the amount representing 100% recovery). Note that the two figures are the same data; however, Figure 3 presents individual swab technician results and Figure 4 presents the aggregate data. Although no conclusion is drawn from the data about the effect of spiked amount on percent recovery, the linear relationships suggests that percent recovery is the same over the evaluated range. For this study, the range from the low spiked level to the high spiked level was a factor of about five (5).
Another publication is by C. Glover (Validation of the Total Organic Carbon (TOC) Swab Sampling and Test Method, Journal of Pharmaceutical Science and Technology, September-October 2006, pp. 284-290). In Table I, percent recovery is given as function of five spiked levels from 5 µg to 100 µg (a range representing a factor of 20). The reported data showed a general decrease in percent recovery as a function of increasing spiked amount. However, the data at the lower levels gave recoveries of greater than 150%, while the recoveries at the higher levels were close to 100%. That data suggests some issues with the TOC analysis.
A final publication is M. A. Strege et al (Total Organic Carbon Analysis of Swab Samples for the Cleaning Validation of Bioprocess Fermentation Equipment, BioPharm International, April 1996). In this publication, Table 2 lists the percent recoveries from stainless steel for three dilutions of a fermentation cell paste. In this situation, the recovery increased with increasing spike level from a low of 75% to a high of 103%. This involved a range with a factor of four between the top and bottom spiked levels.
So, where does this leave us? The data from published studies sometimes show increasing percent recoveries with increasing spiked amounts, sometimes show no change in percent recoveries with increasing spiked amount, and sometimes show decreasing recoveries with increasing spiked amount. I should point out that in none of these studies cited was the stated objective to determine the relationship between the percent recovery and the amount of residue spiked.
If anyone has published any studies that can help elucidate this issue, I’d like to hear from them. If anyone would like to perform a study to specifically evaluate the relationship between percent recovery and spiked amount, I would be more than happy to assist in the design of it so that appropriate conclusions can be drawn. My only caution is that it would be best to avoid using TOC as the analytical method because of control of the sources of TOC. Furthermore, randomization of sampling order must be considered.
Until such time as a definitive study is published and confirmed, it would probably be best to stick with my original contention (based on a common sense understanding of what happens in a swabbing procedure) that other things being equal, percent recovery decreases with increasing spiked level, without stating how significant that decrease might be.
Understanding the Cleaning Process in 2010
In January 2005 I wrote a Cleaning Memo entitled “Understanding the Cleaning Process”. That Cleaning Memo was in response to the FDA report on risk-based approaches to pharmaceutical CGMPs. At the end of that Cleaning Memo, I encouraged manufacturers to “explore more fully what is occurring in a cleaning process, and then to use that knowledge to design more effective and more efficient cleaning processes, as well as simpler ways to validate those processes.” That encouragement is even more important now, based on the 2008 FDA draft process validation guidance (which reportedly will be finalized in the first quarter of 2011) that defines “design and development” as the first stage of the validation process.
Understanding what is happening in the cleaning process (which also includes what is happening in the equipment soiling process) is a key to using these new principles to more effectively and efficiently implement cleaning validation. One of the least quoted sections of the 1993 FDA cleaning validation guidance is in Section IV (“Evaluation of Cleaning Validation”) where the statement is made that “Answers to these questions may also identify steps that can be eliminated for more effective measures and result in resource savings for the company.” This statement, in which the FDA is suggesting “resource saving”, is made in the context of question such as “… at what point does a piece of equipment or system become clean? Does it have to be scrubbed by hand? What is accomplished by hand scrubbing rather than just a solvent wash? How variable are manual cleaning processes from batch to batch and product to product?” Now don’t think for a minute that the FDA wants “resource savings” so that your firm can be more profitable. I don’t know what the FDA’s reason for this statement was in 1993, but in 2010 the reason is clearly that the FDA is concerned about costs of drugs to consumers.
So, when the FDA is encouraging the industry to look for “resource savings”, why are some of us still doing things the way we did in 1993 when we had more limited information of what was happening in our cleaning processes? What are some things that we can do based on a better understanding of the cleaning process?
Well, one example is the ISPE RiskMaPP approach to dealing with highly hazardous drug actives (such as actives that might be genotoxic, mutagenic, or teratogenic). While I criticized (in my November 2010 Cleaning Memo) the RiskMaPP document for the way it critiqued previous methods of setting limits, the fundamental approach in that RiskMaPP document of setting health-based limits for highly hazardous actives (as opposed to previous approaches of using dedicated equipment or requiring that limits be set as non-detectable by the best available analytical technique) is an good example of using knowledge of the cleaning process to enable these highly hazardous actives to be manufactured in the same facility or on the same equipment (with appropriate controls in place).
Another example is understanding what is happening during the dirty hold time (DHT). The traditional approach has been to require a challenge of the maximum time during a validation protocol. However, it is clear (at least in some cases) that an increase in time does not change the difficulty of cleaning. In other cases, the difficulty of cleaning may increase with time up to a certain point (for example, when a liquid is “dry”), and then not change after that. An evaluation of how to approach the DHT will depend on our understanding of how the diffculty of cleaning might change over time. That means not only understanding the nature of drying, but also whether bioburden proliferates during the DHT and whether degradation of the active is accelerated during the DHT. Of course, the approach now should be to understand those factors, and design the cleaning process with those factors in mind, such that the challenges are addressed during the design/development stage rather than during the qualification protocol.
A third example is dealing with campaign length, where so-called minor cleaning (such as vacuuming or a water flush) is performed between batches, and a validated cleaning process is only performed at the end of a campaign. The question comes up, what if the campaign is sometimes 5 batches and sometime 7 batches? Do my qualification protocols have to be at the maximum campaign length? Well, absent any information on the effect of campaign length on the diffculty of cleaning, it makes sense to perform a qualification protocol at the end of the maximum of 7 batches. If only production scheduling would cooperate by providing those number of batches for the required number of validation runs, it might be easy. However, scheduling is not always that nice; plus there might be a time when they want to run 8 batches in a campaign. What can be done?
Well, the secret phrase in the above paragraph was “absent any information on the effect of campaign length on the diffculty of cleaning”. Is there information or data I can obtain from laboratory or developmental studies, or from “sufficiently similar” products or processes, that might allow me to determine the effect of campaign length on diffculty of cleaning? Particularly if I can demonstrate that campaign length has no effect on difficulty of cleaning, performing my qualification protocols after just one batch may be adequate. Again, this may also involve determining information such as bioburden proliferation and/or degradation of the active as a function of campaign length.
There are other approaches to understanding the cleaning process which were discussed in my April 2008 Cleaning Memo (“What Have We Learned in the Last Two Decades?”) that also can be considered. The issue is this – if a pharmaceutical manufacturer is to thrive in the coming decades, the approach of new drugs with significant advantages is still the primary objective. However, low cost (but meeting current CGMP) production should be a secondary, but vital objective.
Understanding what is happening in the cleaning process (which also includes what is happening in the equipment soiling process) is a key to using these new principles to more effectively and efficiently implement cleaning validation. One of the least quoted sections of the 1993 FDA cleaning validation guidance is in Section IV (“Evaluation of Cleaning Validation”) where the statement is made that “Answers to these questions may also identify steps that can be eliminated for more effective measures and result in resource savings for the company.” This statement, in which the FDA is suggesting “resource saving”, is made in the context of question such as “… at what point does a piece of equipment or system become clean? Does it have to be scrubbed by hand? What is accomplished by hand scrubbing rather than just a solvent wash? How variable are manual cleaning processes from batch to batch and product to product?” Now don’t think for a minute that the FDA wants “resource savings” so that your firm can be more profitable. I don’t know what the FDA’s reason for this statement was in 1993, but in 2010 the reason is clearly that the FDA is concerned about costs of drugs to consumers.
So, when the FDA is encouraging the industry to look for “resource savings”, why are some of us still doing things the way we did in 1993 when we had more limited information of what was happening in our cleaning processes? What are some things that we can do based on a better understanding of the cleaning process?
Well, one example is the ISPE RiskMaPP approach to dealing with highly hazardous drug actives (such as actives that might be genotoxic, mutagenic, or teratogenic). While I criticized (in my November 2010 Cleaning Memo) the RiskMaPP document for the way it critiqued previous methods of setting limits, the fundamental approach in that RiskMaPP document of setting health-based limits for highly hazardous actives (as opposed to previous approaches of using dedicated equipment or requiring that limits be set as non-detectable by the best available analytical technique) is an good example of using knowledge of the cleaning process to enable these highly hazardous actives to be manufactured in the same facility or on the same equipment (with appropriate controls in place).
Another example is understanding what is happening during the dirty hold time (DHT). The traditional approach has been to require a challenge of the maximum time during a validation protocol. However, it is clear (at least in some cases) that an increase in time does not change the difficulty of cleaning. In other cases, the difficulty of cleaning may increase with time up to a certain point (for example, when a liquid is “dry”), and then not change after that. An evaluation of how to approach the DHT will depend on our understanding of how the diffculty of cleaning might change over time. That means not only understanding the nature of drying, but also whether bioburden proliferates during the DHT and whether degradation of the active is accelerated during the DHT. Of course, the approach now should be to understand those factors, and design the cleaning process with those factors in mind, such that the challenges are addressed during the design/development stage rather than during the qualification protocol.
A third example is dealing with campaign length, where so-called minor cleaning (such as vacuuming or a water flush) is performed between batches, and a validated cleaning process is only performed at the end of a campaign. The question comes up, what if the campaign is sometimes 5 batches and sometime 7 batches? Do my qualification protocols have to be at the maximum campaign length? Well, absent any information on the effect of campaign length on the diffculty of cleaning, it makes sense to perform a qualification protocol at the end of the maximum of 7 batches. If only production scheduling would cooperate by providing those number of batches for the required number of validation runs, it might be easy. However, scheduling is not always that nice; plus there might be a time when they want to run 8 batches in a campaign. What can be done?
Well, the secret phrase in the above paragraph was “absent any information on the effect of campaign length on the diffculty of cleaning”. Is there information or data I can obtain from laboratory or developmental studies, or from “sufficiently similar” products or processes, that might allow me to determine the effect of campaign length on diffculty of cleaning? Particularly if I can demonstrate that campaign length has no effect on difficulty of cleaning, performing my qualification protocols after just one batch may be adequate. Again, this may also involve determining information such as bioburden proliferation and/or degradation of the active as a function of campaign length.
There are other approaches to understanding the cleaning process which were discussed in my April 2008 Cleaning Memo (“What Have We Learned in the Last Two Decades?”) that also can be considered. The issue is this – if a pharmaceutical manufacturer is to thrive in the coming decades, the approach of new drugs with significant advantages is still the primary objective. However, low cost (but meeting current CGMP) production should be a secondary, but vital objective.
A Critique of Cleaning Validation Issues in ISPE’s RiskMaPP
ISPE has issued the document “Risk-Based Manufacture of Pharmaceutical Products” (Volume 7 of their Baseline Guides, September 2010). As I understand it, the effort to write this guide started with concerns over regulatory bodies tending to require dedicated equipment and/or facilities for certain highly hazardous drug actives, such as potent drugs, hormones, genotoxic compounds, and cytotoxic compounds. The major rationale for this ISPE guide was to counteract this approach by providing for an analysis of safety/toxicity data of these “highly hazardous” drug actives to determine a level that might be a negligible (but acceptable) risk in other drug products (thus allowing, with appropriate controls, the ability to manufacture in non-dedicated equipment/facilities).
This effort involved setting limits for these “highly hazardous” actives based on what is called a “health based limit”, that is a limit based on a toxicological evaluation of the relevant safety data (typically a No Observable Adverse Effect Level or NOAEL). This health-based limit is called ADE, or Acceptable Daily Exposure, in the ISPE guide. This effort for setting limits for cleaning purposes for these highly hazardous actives is to be applauded.
Unfortunately, the guide goes beyond that basic focus to discuss cleaning validation in general, and to imply that this method for cleaning validation is appropriate in all cases, including what I will call “conventional” actives that don’t have these highly hazardous concerns. Specifically, the guide states that current methods of setting limits such as “1/1,000th of the lowest clinical dose or 10 ppm in a batch” are “arbitrary limits”. (page 42) Furthermore, “the use of arbitrary non-health-based limits is not scientifically justified” if data exists to calculate a health-based limit. (page 45) Additionally, “Another non-science-based approach for setting cleaning limits is the use of the ‘10 ppm’ specification” (page 46), and “if default values such as 10 [ppm] are used arbitrarily to set allowable residue limits, they may be lower than they need to be from a health perspective” (page 45). It further states that using such values ignores toxicological data and can be “too restrictive or not sufficiently restrictive”. (page 42)
In other words, according to this guide, current methods of setting limits for cleaning validation purposes (which have been used for at least the last 17 years) are “arbitrary” and “not science-based”. I do not find those assertions to be factually correct, nor is there anything in the guide to support those assertions.
Part of the issue is that the ISPE guide sets up a “straw man” in terms of how limits are currently set. Criteria such as 1/1,000th dose or 10 ppm are critiqued individually. As used by manufacturers, the current method (note that this is what I consider the best approach; I am not suggesting that all pharmaceutical companies use this approach) involves setting limits based on the most stringent of these three criteria:
It is latter concern (of dedicating equipment or establishing limits as non-detectable) that I would think that the ISPE guide would want to counteract by offering the idea of a toxicological evaluation to set an acceptable limit. I don’t see the value (or rationale) for the ISPE guide stating that current methods of setting limits are “arbitrary” and “nonscientific”.
Some of the statements made in the ISPE guide about the current way of setting are true. If limits are set at 1/1,000th of a dose, with some conventional drug actives the level of protection will be more than is necessarily required from a patient safety concern. This is to be expected with a “one size fits all” approach (but don’t get me wrong; this one size fits all doesn’t apply to active where the hazard is not related to the therapeutic effect). However, this does not mean that limits are set in a non-scientific way.
In the training session introducing the guide, one of the speakers stated that there is a degree of judgment in establishing appropriate factors for calculating health-based limits. Does that mean that the limits are non-scientific? Of course not! What it means is that different toxicologists may come up with different numbers for the health-based limits, one being more stringent than the other. An analogous situation exists with the 1/1,000th calculation; it is more stringent than it needs to be for some actives, but for all conventional actives it provides a safe limit.
If the health-based criterion is applied to conventional actives, the acceptable levels given in some example in the ISPE guide are extremely high, such that there might be other concerns other than patient safety. Those concerns might include stability, production efficiency, and interference with the bioavailability of the next active. For example, an ADE of an unspecified NSAID active is given as 40 mg/day. A daily dose of this NSAID is given as 800 mg. (page 101) The implication here is that 1/20th of the daily dose is a safe level to have in a daily dose of a subsequent product. If this ADE were present in a subsequently manufactured drug product involving 250 mg tablets given eight tablets per day, the acceptable concentration of that active in that next product would be 20,000 ppm (2%). Would anyone really consider allowing any drug active to be in a subsequent drug product at that level? It just wouldn’t be CGMP. In case you might think I am overstating the case, there is another example given for an “antisense” active, where the ADE is 0.5 mg/day and the daily dose is 10 mg. (page 102) The implication here is 1/20th of a dose of the antisense active is a safe level. Again, this might be safe from a toxicologist viewpoint, but I doubt if most companies would (or should) allow or permit such levels.
Furthermore, at those high levels, it is likely that residue on equipment surfaces (those that could be observed) would be visually dirty. Does it make sense to set limits to allow such a situation (unless for conventional drugs we really want to only require “visually clean” for surfaces which are readily observable)?
The reply to my objection might be given that I have not read the entire document, and that just because these are limits, it doesn’t mean that manufacturers want to allow residue nearing those levels. The concept of “Margin of Safety” is presented in the ISPE document (page 42). The “Margin of Safety” is not to be confused with so-called “safety factors” that are used in setting limits based on a fraction of the dose. The “Margin of Safety” as defined in the ISPE guide, is the difference or “distance” between the established limit and the actual residue data obtained in a cleaning validation protocol. The argument might be made that while the safe limit (as determined by the ADE) is relatively high, the actual data is much below that ADE-based limit., so the situations I discussed are not likely to happen.
My response to that reply is that those situations are allowed if the limit is that high. Yes, under either ADE limits or 1/1,000th limits, I would like my residue data in protocols to be significantly below my established limit. In other words, I want a robust cleaning process. But this is where the ISPE guide and I differ. The ISPE guide states that “Evaluation of the cleaning validation data is the only way to ensure that any residuals after cleaning are as low as possible below the health-based criteria….” [emphasis added] (page 43) It sounds like the guide, while arguing that limits in some cases are more stringent than they need to be, suggests that manufacturers should still clean to residue levels as low as possible. So that you don’t think I am taking this out of context, the guide also states that “Efforts should be made to ensure that cleaning procedures provide as large a safety margin as possible.” (page 45) In other words, the guide seems to be saying that for conventional actives, limits can be looser (that is, higher), but you should still clean to the same low level so that the “Margin of Safety” is as great as possible.
There are other concerns I have about the document. However, the main concern I have is the way the document inappropriately “trashes” current methods of setting limits as arbitrary and non-scientific. I think an easy fix could be made to the document by revising it to what I think is its original focus. That is, take out all general references to how limits are set, and focus on the main issue, which is that health-based limits are appropriate for dealing with drug actives that have considerable hazard concerns unrelated to the therapeutic effect. Those actives include hormones, potent steroids, and actives that are genotoxicity, cytotoxicity, or have reproductive hazards. The rationale for this suggested change is that the purpose of the ISPE guide seems to be having a science- based method for avoiding manufacture of such compounds in dedicated facilities/equipment. Supposedly there is in development another ISPE guide on cleaning validation. It would seem that that cleaning validation guide should be the one to specifically address the issue of whether limits for “conventional” actives should be changed.
Some of you might want to know why I am voicing my objections now. After all, the guide has been in development for five years. In the fall of 2007, I contacted ISPE about obtaining the draft document to provide comments, but found out that the comment period had just closed. However, I voiced my objections (essentially the same objections expressed here) back in 2008. These concerns were not specifically about the draft of the document at that time, but were based on presentations made by Andy Walsh (a RiskMaPP task force member) at an ISPE meeting in Washington DC in the summer of 2008. I gave a webinar entitled “Are we Setting Limits Correctly?” in August 2008 and wrote a Cleaning Memo with the same title in October 2008. However, whatever happened in the past, it should be clear that the published guide needs correction.
For clarification, I am not concerned about the guide’s approach to setting limits for highly hazardous actives. The approach of a toxicological evaluation based on those hazards is appropriate. If those evaluations result in making products in non-dedicated facilities/equipment, then that is a needed step forward. Note that this approach is consistent with the December 2009 EMA document EMA/INS/GMP/809387/2009,“Update on revision of Chapters 3 and 5 of the GMP Guide: Dedicated facilities"
This effort involved setting limits for these “highly hazardous” actives based on what is called a “health based limit”, that is a limit based on a toxicological evaluation of the relevant safety data (typically a No Observable Adverse Effect Level or NOAEL). This health-based limit is called ADE, or Acceptable Daily Exposure, in the ISPE guide. This effort for setting limits for cleaning purposes for these highly hazardous actives is to be applauded.
Unfortunately, the guide goes beyond that basic focus to discuss cleaning validation in general, and to imply that this method for cleaning validation is appropriate in all cases, including what I will call “conventional” actives that don’t have these highly hazardous concerns. Specifically, the guide states that current methods of setting limits such as “1/1,000th of the lowest clinical dose or 10 ppm in a batch” are “arbitrary limits”. (page 42) Furthermore, “the use of arbitrary non-health-based limits is not scientifically justified” if data exists to calculate a health-based limit. (page 45) Additionally, “Another non-science-based approach for setting cleaning limits is the use of the ‘10 ppm’ specification” (page 46), and “if default values such as 10 [ppm] are used arbitrarily to set allowable residue limits, they may be lower than they need to be from a health perspective” (page 45). It further states that using such values ignores toxicological data and can be “too restrictive or not sufficiently restrictive”. (page 42)
In other words, according to this guide, current methods of setting limits for cleaning validation purposes (which have been used for at least the last 17 years) are “arbitrary” and “not science-based”. I do not find those assertions to be factually correct, nor is there anything in the guide to support those assertions.
Part of the issue is that the ISPE guide sets up a “straw man” in terms of how limits are currently set. Criteria such as 1/1,000th dose or 10 ppm are critiqued individually. As used by manufacturers, the current method (note that this is what I consider the best approach; I am not suggesting that all pharmaceutical companies use this approach) involves setting limits based on the most stringent of these three criteria:
- 1/1,000th of a dose of an active in a dose of the next drug product
- 10 ppm of the active in a dose of the next drug product
- Equipment surfaces are visually clean
It is latter concern (of dedicating equipment or establishing limits as non-detectable) that I would think that the ISPE guide would want to counteract by offering the idea of a toxicological evaluation to set an acceptable limit. I don’t see the value (or rationale) for the ISPE guide stating that current methods of setting limits are “arbitrary” and “nonscientific”.
Some of the statements made in the ISPE guide about the current way of setting are true. If limits are set at 1/1,000th of a dose, with some conventional drug actives the level of protection will be more than is necessarily required from a patient safety concern. This is to be expected with a “one size fits all” approach (but don’t get me wrong; this one size fits all doesn’t apply to active where the hazard is not related to the therapeutic effect). However, this does not mean that limits are set in a non-scientific way.
In the training session introducing the guide, one of the speakers stated that there is a degree of judgment in establishing appropriate factors for calculating health-based limits. Does that mean that the limits are non-scientific? Of course not! What it means is that different toxicologists may come up with different numbers for the health-based limits, one being more stringent than the other. An analogous situation exists with the 1/1,000th calculation; it is more stringent than it needs to be for some actives, but for all conventional actives it provides a safe limit.
If the health-based criterion is applied to conventional actives, the acceptable levels given in some example in the ISPE guide are extremely high, such that there might be other concerns other than patient safety. Those concerns might include stability, production efficiency, and interference with the bioavailability of the next active. For example, an ADE of an unspecified NSAID active is given as 40 mg/day. A daily dose of this NSAID is given as 800 mg. (page 101) The implication here is that 1/20th of the daily dose is a safe level to have in a daily dose of a subsequent product. If this ADE were present in a subsequently manufactured drug product involving 250 mg tablets given eight tablets per day, the acceptable concentration of that active in that next product would be 20,000 ppm (2%). Would anyone really consider allowing any drug active to be in a subsequent drug product at that level? It just wouldn’t be CGMP. In case you might think I am overstating the case, there is another example given for an “antisense” active, where the ADE is 0.5 mg/day and the daily dose is 10 mg. (page 102) The implication here is 1/20th of a dose of the antisense active is a safe level. Again, this might be safe from a toxicologist viewpoint, but I doubt if most companies would (or should) allow or permit such levels.
Furthermore, at those high levels, it is likely that residue on equipment surfaces (those that could be observed) would be visually dirty. Does it make sense to set limits to allow such a situation (unless for conventional drugs we really want to only require “visually clean” for surfaces which are readily observable)?
The reply to my objection might be given that I have not read the entire document, and that just because these are limits, it doesn’t mean that manufacturers want to allow residue nearing those levels. The concept of “Margin of Safety” is presented in the ISPE document (page 42). The “Margin of Safety” is not to be confused with so-called “safety factors” that are used in setting limits based on a fraction of the dose. The “Margin of Safety” as defined in the ISPE guide, is the difference or “distance” between the established limit and the actual residue data obtained in a cleaning validation protocol. The argument might be made that while the safe limit (as determined by the ADE) is relatively high, the actual data is much below that ADE-based limit., so the situations I discussed are not likely to happen.
My response to that reply is that those situations are allowed if the limit is that high. Yes, under either ADE limits or 1/1,000th limits, I would like my residue data in protocols to be significantly below my established limit. In other words, I want a robust cleaning process. But this is where the ISPE guide and I differ. The ISPE guide states that “Evaluation of the cleaning validation data is the only way to ensure that any residuals after cleaning are as low as possible below the health-based criteria….” [emphasis added] (page 43) It sounds like the guide, while arguing that limits in some cases are more stringent than they need to be, suggests that manufacturers should still clean to residue levels as low as possible. So that you don’t think I am taking this out of context, the guide also states that “Efforts should be made to ensure that cleaning procedures provide as large a safety margin as possible.” (page 45) In other words, the guide seems to be saying that for conventional actives, limits can be looser (that is, higher), but you should still clean to the same low level so that the “Margin of Safety” is as great as possible.
There are other concerns I have about the document. However, the main concern I have is the way the document inappropriately “trashes” current methods of setting limits as arbitrary and non-scientific. I think an easy fix could be made to the document by revising it to what I think is its original focus. That is, take out all general references to how limits are set, and focus on the main issue, which is that health-based limits are appropriate for dealing with drug actives that have considerable hazard concerns unrelated to the therapeutic effect. Those actives include hormones, potent steroids, and actives that are genotoxicity, cytotoxicity, or have reproductive hazards. The rationale for this suggested change is that the purpose of the ISPE guide seems to be having a science- based method for avoiding manufacture of such compounds in dedicated facilities/equipment. Supposedly there is in development another ISPE guide on cleaning validation. It would seem that that cleaning validation guide should be the one to specifically address the issue of whether limits for “conventional” actives should be changed.
Some of you might want to know why I am voicing my objections now. After all, the guide has been in development for five years. In the fall of 2007, I contacted ISPE about obtaining the draft document to provide comments, but found out that the comment period had just closed. However, I voiced my objections (essentially the same objections expressed here) back in 2008. These concerns were not specifically about the draft of the document at that time, but were based on presentations made by Andy Walsh (a RiskMaPP task force member) at an ISPE meeting in Washington DC in the summer of 2008. I gave a webinar entitled “Are we Setting Limits Correctly?” in August 2008 and wrote a Cleaning Memo with the same title in October 2008. However, whatever happened in the past, it should be clear that the published guide needs correction.
For clarification, I am not concerned about the guide’s approach to setting limits for highly hazardous actives. The approach of a toxicological evaluation based on those hazards is appropriate. If those evaluations result in making products in non-dedicated facilities/equipment, then that is a needed step forward. Note that this approach is consistent with the December 2009 EMA document EMA/INS/GMP/809387/2009,“Update on revision of Chapters 3 and 5 of the GMP Guide: Dedicated facilities"
Plasmids for Vaccine Validation
pCMV-S (also known as pRc/CMV-HBs) (Figure 5) is widely used to validate DNA vaccine delivery and formulation strategies. This plasmid expresses the hepatitis B surface antigen (HBsAg) under the control of the CMV immediate-early promoter.
pCMVHB-S2.S expresses the small and middle forms of recombinant HBsAg. The plasmid can be used to generate anti-HBsAg antibodies like pCMV-S. It is also used to fuse other sequences to the S form of HBsAg.
The antibody response to pCMV-S and/or pCMVHB-S2.S can be tested via ELISA using Aldevron's recombinant HBsAg. Both of these plasmids are available free of charge for research applications (Table 5). These vectors are covered and described by United States Patent 6,635,624 which is available at www.uspto.gov.
pCMVHB-S2.S expresses the small and middle forms of recombinant HBsAg. The plasmid can be used to generate anti-HBsAg antibodies like pCMV-S. It is also used to fuse other sequences to the S form of HBsAg.
The antibody response to pCMV-S and/or pCMVHB-S2.S can be tested via ELISA using Aldevron's recombinant HBsAg. Both of these plasmids are available free of charge for research applications (Table 5). These vectors are covered and described by United States Patent 6,635,624 which is available at www.uspto.gov.
Figure 5 — Validate your DNA vaccine delivery and formulation with plasmids provided free, courtesy of Aldevron.
Subscribe to:
Posts (Atom)